Pith. sign in

REVIEW 4 major objections 6 minor 61 references

Can Artificial Intelligence Write Like Borges? An Evaluation Protocol for Spanish Microfiction

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper claims that a fifteen-question protocol grounded in editorial practice can reliably assess the literary value of human-written and AI-generated Spanish microfictions, with expert ratings that track author experience.

desk verdict A promising literary-theory-based evaluation protocol for Spanish microfiction, but the reliability validation is confounded by a two-arm design where experts see only human texts and enthusiasts only AI texts. read the letter →

arxiv 2506.08172 v1 pith:24GMWODM submitted 2025-06-09 cs.CL

classification cs.CL
keywords artificialintelligencecreativewritingevaluationprotocolmicrofictionGrAImesliteraryreceptionlargelanguagemodelsinter-raterreliability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the literary quality of microfiction can be scored with a reusable instrument rather than by intuition or surface text metrics. The instrument, GrAImes, is a fifteen-item questionnaire — ten Likert-scale items and five open answers — spanning literary interpretation, technical craft, and editorial or commercial appeal, modeled on how publishers decide whether to accept a manuscript. The authors validate it in two small experiments: five literature experts rated six human-written microfictions, and sixteen reading enthusiasts rated six AI-generated microfictions, three from ChatGPT-3.5 and three from a GPT-2 model fine-tuned on Spanish microfiction. They report good to acceptable internal consistency across raters, expert ratings that tracked the authors' experience level, and enthusiast ratings that slightly favored ChatGPT-3.5 on commercial appeal. If the protocol holds, it would give computational creativity research a literary-theory-grounded way to compare human and machine writing, and a counterweight to crowd-based claims that AI poetry already outranks the canon.

What carries the argument

The load-bearing object is the GrAImes questionnaire: fifteen items — ten answered on a 1-to-5 Likert scale and five as open answers — split into three dimensions: story overview and text complexity (thematic coherence, clarity, interpretive depth), technical assessment (credibility, reader cooperation, originality of reality, genre, and language), and editorial or commercial quality (intertextual familiarity, desire for more, recommendation, gift-worthiness, publisher fit). Its design mirrors the editorial report a publisher commissions on unsolicited manuscripts, converting the editor's parameters — content clarity, technical value, and relevance — into scored items. The argument is carried by three reliability statistics applied to the raters' responses: the intraclass correlation coefficient for agreement, Cronbach's alpha for the internal consistency of each text's scores, and Kendall's W as a concordance measure chosen because it is less affected by small sample sizes. The protocol's claim to objectivity lives in the inter-rater agreement these statistics report, and its claim to literary validity lives in the questions themselves, which are grounded in reception theory and editorial practice rather than corpus statistics.

What would settle it

Run the same protocol with a larger panel (on the order of thirty experts and a matched set of texts at each experience level); if Cronbach's alpha for expert-authored texts no longer separates from emerging-writer texts, or the ICC values for the Likert items fall below acceptable thresholds, the claimed reliability collapses. A second decisive check is test-retest: have the same raters score the same microfictions twice, weeks apart; if individual ratings drift substantially, the instrument measures transient preference rather than stable literary judgment. A third check is a control panel of readers with no literary training rating the human-written texts; if their scores track the experts', the instrument is registering generic fluency rather than literary quality.

Watch

Extended reading notes

Core claim

GrAImes is presented as a reliable framework for assessing the literary value of human-written and AI-generated microfictions, something the authors argue current natural-language metrics cannot do because BLEU, ROUGE, and perplexity measure surface similarity rather than metaphor, symbolism, or stylistic originality. The protocol turns the publishing industry's editorial report into fifteen questions organized in three dimensions — story overview and textual complexity, technical assessment, and editorial or commercial quality — so that literary value is operationalized as thematic coherence, interpretive depth, technical execution, and market viability. Validation rests on three inter-rater statistics: the intraclass correlation coefficient, Cronbach's alpha, and Kendall's W, applied to two experiments with five expert and sixteen enthusiast raters. The authors report good to acceptable internal consistency, a correlation between author expertise and scores from the expert panel, and a slight enthusiast preference for ChatGPT-3.5 microfictions over the fine-tuned baseline in editorial and commercial appeal, while the baseline scored slightly higher on technical quality. They position GrAImes as a challenge to crowd-sourced findings that non-experts prefer AI-generated poetry to canonical human poetry, arguing that evaluations by readers without literary training measure immediate readability rather than interpretive depth.

Load-bearing premise

The validation rests on two linked premises: that fifteen questions about interpretation, technique, and marketability capture what makes a microfiction literary, and that five experts and sixteen enthusiasts rating six texts each is a large enough sample for the reliability statistics to mean anything.

Editorial extensions

If this is right

  • GrAImes gives researchers a single fifteen-item instrument for comparing human-written, AI-generated, and AI-assisted microfictions on the same literary criteria, replacing surface metrics like BLEU and perplexity for this genre.
  • Because expert ratings tracked author experience (Cronbach's alpha of 0.80 and 0.79 for expert-authored texts versus 0.34 and 0.13 for emerging writers), the protocol can separate more accomplished writing from less accomplished writing.
  • On the enthusiast panel, ChatGPT-3.5 microfictions scored higher on editorial and commercial appeal while the fine-tuned baseline scored slightly higher on technical quality, implying that general audiences reward fluency and marketability more than structural craft.
  • Evaluator composition changes the result: experts emphasized originality and technical execution, enthusiasts emphasized accessibility, so any comparison of human versus AI literary output must report who did the judging.
  • The protocol is positioned as a check on crowd-sourced claims that AI poetry outranks canonical human poetry, by measuring interpretive depth rather than immediate preference.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: A test-retest study — the same evaluators re-rating the same microfictions weeks apart — would separate stable literary judgments from familiarity or mood effects; the paper reports inter-rater agreement but no intra-rater stability, so its reliability claim covers only one axis of reliability.
  • Inference: The instrument's diagnostic profile (high technical scores, low innovation scores across both panels) suggests it could steer generation: fine-tuning or prompting could target the weakest dimensions, such as proposing a new vision of the genre, which scored lowest almost everywhere.
  • Inference: The negative ICC values for one item in each panel — Question 13 on gift-worthiness at -0.72 among experts and Question 8 on genre innovation at -0.44 among enthusiasts — indicate that some items actively depress the summary reliability figures; a revision that rewrites or drops these items would likely change the headline 'good to acceptable' verdict.
  • Inference: Applying GrAImes beyond Spanish would require re-norming rather than translation alone, since several items (publisher fit, gift-worthiness, intertextual recognition) presuppose a specific literary market and readership; otherwise cross-linguistic comparisons would confound literary quality with cultural familiarity.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces GrAImes, a 15-item evaluation protocol for Spanish microfictions, grounded in literary theory and editorial practice, and reports two validation experiments. In the first, five PhD-holding literary experts rated six human-authored microfictions; in the second, sixteen literature enthusiasts rated six AI-generated microfictions (three from ChatGPT-3.5 and three from Monterroso, a GPT-2 baseline fine-tuned on Spanish microfiction). The authors report internal-consistency statistics (ICC, Cronbach's alpha, Kendall's W), a Sentence-BERT comparison of open-answer responses, and expert feedback on the protocol. They conclude that GrAImes 'could become a reliable framework' for assessing literary quality and that ChatGPT-3.5 texts were slightly favored over Monterroso texts. The paper includes a GitHub repository for reproducibility.

Significance. The paper addresses a genuine gap: automated metrics such as BLEU, ROUGE, and perplexity are not designed to capture literary qualities, and the protocol's grounding in reception theory and editorial criteria is a welcome contribution. The use of real PhD-level literary experts, the inclusion of a fine-tuned Spanish microfiction baseline, and the public GitHub repository for replication are strengths. If the validation were properly designed, GrAImes could be a useful instrument for the computational-creativity community. However, the current evidence does not support the broad reliability claim because the experimental design confounds rater expertise with text source, and the sample sizes are too small for the reported reliability statistics to be stable.

major comments (4)
  1. [Sections 3.2.1-3.2.2 and 4] The validation design is confounded between rater group and text source: the five experts evaluated only the six human-written microfictions, and the sixteen enthusiasts evaluated only the six AI-generated microfictions. The paper's central claim that GrAImes has 'good to acceptable internal consistency' for assessing both human and AI-authored texts (Abstract; Section 4) is therefore not established, because reliability could depend on rater expertise, text type, or their interaction, and the design contains no Expert×AI or Enthusiast×Human cell. The authors should either cross the design or explicitly restrict the reliability claim to the measured cells.
  2. [Section 3.3, Tables 6-7 and 13] The reliability statistics are computed on very small samples: Cronbach's alpha is estimated per microfiction with only five expert raters (Table 7) and sixteen enthusiast raters (Table 13), and the reported point values (e.g., 0.80, 0.79) have wide confidence intervals that include unacceptable levels of consistency. The paper acknowledges sample-size sensitivity in Section 3.3 but still treats these point estimates as confirmatory; the authors should report confidence intervals or bootstrap estimates and temper the conclusions accordingly.
  3. [Section 4.1, Tables 3-7] The labeling of the human-authored microfictions is internally inconsistent: Table 3 assigns MF3 and MF6 to the 'Medium' experience author, but the text in Section 4.1 describes MF3 and MF5 as written by 'low expertise' authors, describes MF6 as by an 'emerging author', and repeats a contradictory sentence about MF4 and MF6 versus MF3 and MF6. These contradictions undermine the reported correlation between author expertise and expert evaluations and make the results non-reproducible; the authors must correct the mismatches between the table and the prose.
  4. [Section 4.2, Tables 14-16] Tables 14-16 are headed 'Literary experts' responses' to Monterroso and ChatGPT-3.5 microfictions, which contradicts Section 3.2.2 where the AI-generated texts are assigned to the Enthusiast group. The surrounding text alternates between 'Enthusiast group leaders' and 'literary experts', making it unclear which rater population produced these data. In addition, the comparison between ChatGPT-3.5 and Monterroso is based on descriptive averages only, with no significance test; the claim that ChatGPT texts were 'slightly favored' is not statistically supported. The authors should clarify the rater groups and provide appropriate inferential statistics.
minor comments (6)
  1. [Section 4.1] The section title and text contain several typos: 'GrAlmes' instead of 'GrAImes' and 'Cronbanch' instead of 'Cronbach' in Table 7.
  2. [Figure 10 and Figure 14] The captions contain typos: 'nthusiast' in Figure 10 and 'evlauation' in Figure 14 should be corrected.
  3. [Section 2.2] The sentence 'a critical stand is needeed' contains a typo and should read 'needed'.
  4. [Table 8] The entry for MF3, Question 6 reads '4.3 1.' with an incomplete decimal; this should be corrected to a complete value.
  5. [Section 4.2] The phrase 'MF 3, generated by Program A' should read 'generated by Monterroso' for consistency with the rest of the section.
  6. [Section 3.2.2] The description of the Enthusiast group mentions '16 literary enthousiasts, plus the group leader and booktuber', but the results in Section 4.2 sometimes refer to 'group leaders' without specifying whether the leader is included in the reported statistics; this ambiguity should be resolved.

Circularity Check

0 steps flagged · score 2.0 of 10

No derivational circularity; the GrAImes reliability claim rests on independent rating data, and the paper contains only a minor, non-load-bearing self-citation.

full rationale

The derivation chain is not circular. GrAImes is a 15-item questionnaire constructed from literary-theory and editorial-practice considerations (Section 3.1, Table 2), not fitted to the data it later evaluates. The reliability statistics (ICC, Cronbach's alpha, Kendall's W; Section 3.2.4) are computed on evaluators' ratings and are not quantities defined in terms of their own outputs; no predicted value is forced by a fitted parameter. The only self-referential elements are the authors' own prior genre-definition work [18] used in Section 2.1, an accompanying microfiction by an author (Figure 1), and the expert poll in Section 4.1 asking whether the protocol works ('A strong consensus (4 out of 5 experts) agreed that the protocol can effectively evaluate the literary value of microfiction'). The poll is face-validity evidence, not a derived prediction, and the self-citation is descriptive and corroborated by external sources [17,21], so neither is load-bearing. The main weaknesses are methodological: expert/enthusiast rater groups are perfectly confounded with human/AI text sources, and per-text Cronbach's alpha with five raters is statistically unstable. These limitations weaken the reliability inference, but confounding and small samples are not circularity. Score 2 reflects the minor non-load-bearing self-citation, not a reduction of the central claim to its inputs.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The protocol relies on literary-theory assumptions about what makes a text literary, a genre definition, and on standard reliability statistics. There are no fitted numerical parameters and no newly postulated physical or conceptual entities beyond the questionnaire itself.

assumptions (4)
  • domain assumption Literary texts are characterized by verisimilitude, codification, rule-breaking, and deferred communication (Section 3.1).
    These four features, drawn from pragmatic and literary theory, are the basis for the GrAImes questionnaire items.
  • domain assumption Reception theory: meaning is co-produced by readers, whose literary competence shapes their evaluation (Sections 1 and 2.2).
    The paper uses this to justify selecting experts and enthusiasts as evaluators rather than crowd-sourced general readers.
  • domain assumption Microfiction is defined as a narrative text of up to 300 words (Section 2.1).
    The word limit sets the genre boundary for the experiments; the paper adopts Ana Maria Shua's definition.
  • standard math ICC, Cronbach's alpha, and Kendall's W are appropriate reliability measures for the sample sizes used (Section 3.2.4).
    The paper explicitly acknowledges ICC and Cronbach's alpha are sensitive to small samples and uses Kendall's W as a more robust alternative; the validity of the reliability conclusions depends on these standard statistical models.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can Artificial Intelligence Write Like Borges? An Evaluation Protocol for Spanish Microfiction." pith.science (2026). https://pith.science/paper/24GMWODM

@misc{pith2026250608172,
  author       = {Pith},
  title        = {Pith review of: Can Artificial Intelligence Write Like Borges? An Evaluation Protocol for Spanish Microfiction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/24GMWODM}},
  note         = {Machine review of arXiv:2506.08172}
}
read the original abstract

Automated story writing has been a subject of study for over 60 years. Large language models can generate narratively consistent and linguistically coherent short fiction texts. Despite these advancements, rigorous assessment of such outputs for literary merit - especially concerning aesthetic qualities - has received scant attention. In this paper, we address the challenge of evaluating AI-generated microfictions and argue that this task requires consideration of literary criteria across various aspects of the text, such as thematic coherence, textual clarity, interpretive depth, and aesthetic quality. To facilitate this, we present GrAImes: an evaluation protocol grounded in literary theory, specifically drawing from a literary perspective, to offer an objective framework for assessing AI-generated microfiction. Furthermore, we report the results of our validation of the evaluation protocol, as answered by both literature experts and literary enthusiasts. This protocol will serve as a foundation for evaluating automatically generated microfictions and assessing their literary value.

Figures

Figures reproduced from arXiv: 2506.08172 by the authors.

Figure 1
Figure 1. This figure presents an example of a microfiction authored by Yobany Garcia Medina. It is divided into three parts: opening (A), development (B), and closing (C), collectively forming a cohesive narrative unit. Reading microfiction requires the reader not only to interpret its meaning but also to reconstruct its structure. The text prompts the reader to complete the narrative by providing cues that suggest a storyli… view at source ↗
Figure 2
Figure 2. Microfiction example [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Microfiction evaluation process applied GrAImes to stories generated by language models, evaluated by a community of reading and literature enthusiasts. 3.2.1. Evaluators We gathered two groups of evaluators: the Experts and the Enthusiasts. The selection criteria for the Expert group consisted of five literary scholars, each holding a PhD in Spanish or Latin American literature and occupying a permanent academic po… view at source ↗
Figures from the paper (13 more)
Figure 4
Figure 4. Figure 4: Expert group evaluation averages where: n is the number of objects, m is the number of judges and S is the sum of ranks squared deviations. In this study, we aimed to evaluate the literary quality of Spanish-language microfiction, focusing on both human and AI-generate…
Figure 5
Figure 5. Figure 5: Expert group evaluation of human written microfictions using GrAImes (averages by section) From the responses obtained and displayed in Tables 4 and 5, we conclude that literary experts rated the microfictions (1 and 2) authored by an expert writer more favorably. Howe…
Figure 6
Figure 6. Figure 6: ICC and Cronbach’s Alpha line charts of the Expert group evaluation of human written microfictions [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Kendall W by GrAImes sections of the Expert group evaluation of human written microfictions. These findings align with existing research [59] on the relationship between writing expertise and text coherence. Higher expertise leads to better-structured, logically consis…
Figure 8
Figure 8. Figure 8: Interrater agreement on clarity (q1), structure(q2), complexity (q4), giftability (q14), and commerciality (q15) for microfiction 2 from the Expert group On the two extra questions given to the literary experts (see 3.2.2), the majority of experts (3 out of 5) found th…
Figure 9
Figure 9. Figure 9: Enthusiast group evaluation of AI-generated microfictions study evaluated the microfictions based on parameters such as coherence, thematic depth, stylistic originality, and emotional resonance. A total of six microfictions were generated, with three created by the Mon…
Figure 10
Figure 10. Figure 10: nthusiast group evaluation of AI-generated microfictions by section [PITH_FULL_IMAGE:figures/full_fig_p016_10.png]
Figure 11
Figure 11. Figure 11: Line charts of literature enthusiasts GrAImes sections summarized AV and SD that, regardless of other literary attributes, the microfictions maintain a sense of realism that resonates with readers. The question regarding whether the text requires the reader’s particip…
Figure 12
Figure 12. Figure 12: ICC line chart, literature enthusiasts responses to AI generated microfictions. responses, suggesting that some readers found deeper layers of meaning, while others perceived the texts as more straightforward. The Intraclass Correlation Coefficient analysis of GrAImes…
Figure 13
Figure 13. Figure 13: Enthusiast group evaluation of AI-generated microfictions [PITH_FULL_IMAGE:figures/full_fig_p019_13.png]
Figure 14
Figure 14. Figure 14: Enthusiast group evlauation of AI-generated microfictions (average by section) [PITH_FULL_IMAGE:figures/full_fig_p019_14.png]
Figure 15
Figure 15. Figure 15: Line charts of literary experts GrAImes sections summarized AV and SD [PITH_FULL_IMAGE:figures/full_fig_p020_15.png]
Figure 16
Figure 16. Figure 16: Comparison of literary experts vs enthusiasts. GrAImes Sections AV with AI generated MFs. evaluation of microfiction across these two distinct origins—human authors and generative AI. This gap in the literature underscores the significance of our research, which seeks…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

61 extracted references · 53 canonical work pages

  1. [1]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning

    Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P .; Bi, X.; et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948 2025

  2. [2]

    Openai o1 system card

    Jaech, A.; Kalai, A.; Lerer, A.; Richardson, A.; El-Kishky, A.; Low, A.; Helyar, A.; Madry, A.; Beutel, A.; Carney, A.; et al. Openai o1 system card. arXiv preprint arXiv:2412.16720 2024

  3. [3]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

    Team, G.; Georgiev, P .; Lei, V .I.; Burnell, R.; Bai, L.; Gulati, A.; Tanzer, G.; Vincent, D.; Pan, Z.; Wang, S.; et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530 2024

  4. [4]

    AI-generated poetry is indistinguishable from human-written poetry and is rated more favorably

    Porter, B.; Machery, E. AI-generated poetry is indistinguishable from human-written poetry and is rated more favorably. Scientific Reports 2024, 14, 26133

  5. [5]

    Distilling the knowledge in a neural network

    Hinton, G.; Vinyals, O.; Dean, J. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531 2015

  6. [6]

    ‘Frontier AI,’Power, and the Public Interest: Who benefits, who decides? Harvard Data Science Review 2024

    Leslie, D.; Ashurst, C.; González, N.M.; Griffiths, F.; Jayadeva, S.; Jorgensen, M.; Katell, M.; Krishna, S.; Kwiatkowski, D.; Martins, C.I.; et al. ‘Frontier AI,’Power, and the Public Interest: Who benefits, who decides? Harvard Data Science Review 2024

  7. [7]

    Neural text generation in stories using entity representations as context

    Clark, E.; Ji, Y.; Smith, N.A. Neural text generation in stories using entity representations as context. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), 2018, pp. 2250–2260

  8. [8]

    Human heuristics for AI-generated language are flawed

    Jakesch, M.; Hancock, J.T.; Naaman, M. Human heuristics for AI-generated language are flawed. Proceedings of the National Academy of Sciences 2023, 120

Show all 61 references
  1. [9]

    Automatic story generation: a survey of approaches

    Alhussain, A.I.; Azmi, A.M. Automatic story generation: a survey of approaches. ACM Computing Surveys (CSUR) 2021, 54, 1–38

  2. [10]

    The act of reading: A theory of aesthetic response

    Iser, W. The act of reading: A theory of aesthetic response. Journal of Aesthetics and Art Criticism 1979, 38. 3 Jorge Luis Borges (1899–1986) was an Argentine writer, poet, and essayist widely regarded as one of the most influential literary figures of the 20th century. Known...

  3. [11]

    Concretización y reconstrucción

    Ingarden, R. Concretización y reconstrucción. En busca del texto: teoría de la recepción literaria 1993, pp. 31–54

  4. [12]

    Grimes’ Fairy Tales: A 1960s Story Generator; Springer International Publishing: Cham, 2017; pp

    Ryan, J. Grimes’ Fairy Tales: A 1960s Story Generator; Springer International Publishing: Cham, 2017; pp. 89–103

  5. [13]

    The Russian Folktale by Vladimir Yakovlevich Propp; Wayne State University Press, 2012

    Propp, V .Y. The Russian Folktale by Vladimir Yakovlevich Propp; Wayne State University Press, 2012

  6. [14]

    What editors do: the art, craft, and business of book editing; University of Chicago Press, 2017

    Ginna, P . What editors do: the art, craft, and business of book editing; University of Chicago Press, 2017

  7. [15]

    Evaluation of automatic generation of basic stories

    Peinado, F.; Gervás, P . Evaluation of automatic generation of basic stories. New Generation Computing 2006, 24, 289–302

  8. [16]

    The creative mind: Myths and mechanisms; Routledge, 2004

    Boden, M.A. The creative mind: Myths and mechanisms; Routledge, 2004

  9. [17]

    La minificción como clase textual transgenérica

    Tomassini, G.; Maris, S. La minificción como clase textual transgenérica. Revista interamericana de bibliografía: Review of interamerican bibliography 1996, 46, 6–6

  10. [18]

    Microrrelato o minificción: de la nomenclatura a la estructura de un género literario

    Medina, Y.d.J.G. Microrrelato o minificción: de la nomenclatura a la estructura de un género literario. Microtextualidades. Revista Internacional de microrrelato y minificción 2017, pp. 89–102

  11. [19]

    La función narrativa

    Ricoeur, P . La función narrativa. Revista de Semiótica 1989, 1, 69–90

  12. [20]

    La aventura semiológica; Paidós Barcelona, 1990

    Barthes, R.; Alcalde, R. La aventura semiológica; Paidós Barcelona, 1990

  13. [21]

    Cómo escribir un microrrelato; Siglo XXI Editores, 2023

    Shua, A.M. Cómo escribir un microrrelato; Siglo XXI Editores, 2023

  14. [22]

    Hierarchical neural story generation

    Fan, A.; Lewis, M.; Dauphin, Y. Hierarchical neural story generation. arXiv preprint arXiv:1805.04833 2018

  15. [23]

    Long text generation by modeling sentence-level and discourse-level coherence

    Guan, J.; Mao, X.; Fan, C.; Liu, Z.; Ding, W.; Huang, M. Long text generation by modeling sentence-level and discourse-level coherence. arXiv preprint arXiv:2105.08963 2021

  16. [24]

    Deep Learning-Based Short Story Generation for an Image Using the Encoder-Decoder Structure

    Min, K.; Dang, M.; Moon, H. Deep Learning-Based Short Story Generation for an Image Using the Encoder-Decoder Structure. IEEE Access 2021, 9, 113550–113557

  17. [25]

    GPoeT-2: A GPT-2 Based Poem Generator

    Lo, K.L.; Ariss, R.; Kurz, P . GPoeT-2: A GPT-2 Based Poem Generator. arXiv preprint arXiv:2205.08847 2022

  18. [26]

    Character-based interactive storytelling

    Cavazza, M.; Charles, F.; Mead, S.J. Character-based interactive storytelling. IEEE Intelligent systems 2002, 17, 17–24

  19. [27]

    Story plot generation based on CBR.Journal of Knowledge-Based Systems 2005, 18, 235–242

    Gervás, P .; Díaz-Agudo, B.; Peinado, F.; Hervás, R. Story plot generation based on CBR.Journal of Knowledge-Based Systems 2005, 18, 235–242

  20. [28]

    Toward a Better Story End: Collecting Human Evaluation with Reasons

    Mori, Y.; Yamane, H.; Mukuta, Y.; Harada, T. Toward a Better Story End: Collecting Human Evaluation with Reasons. In Proceedings of the 12th International Conference on Natural Language Generation, 2019, pp. 383–390

  21. [29]

    Generating different story tellings from semantic representations of narrative

    Rishes, E.; Lukin, S.M.; Elson, D.K.; Walker, M.A. Generating different story tellings from semantic representations of narrative. In Proceedings of the International Conference on Interactive Digital Storytelling. Springer, 2013, pp. 192–204

  22. [30]

    A Tool for Deep Semantic Encoding of Narrative Texts

    Elson, D.K.; McKeown, K.R. A Tool for Deep Semantic Encoding of Narrative Texts. ACL-IJCNLP 2009, p. 9

  23. [31]

    Generating text with recurrent neural networks; 2011

    Sutskever, I.; Martens, J.; Hinton, G.E. Generating text with recurrent neural networks; 2011

  24. [32]

    Globally coherent text generation with neural checklist models

    Kiddon, C.; Zettlemoyer, L.; Choi, Y. Globally coherent text generation with neural checklist models. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 2016, pp. 329–339

  25. [33]

    Aligning books and movies: Towards story-like visual explanations by watching movies and reading books

    Zhu, Y.; Kiros, R.; Zemel, R.; Salakhutdinov, R.; Urtasun, R.; Torralba, A.; Fidler, S. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision, 2015, pp. 19–27

  26. [34]

    Perceived or not perceived: Film character models for expressive nlg

    Walker, M.A.; Grant, R.; Sawyer, J.; Lin, G.I.; Wardrip-Fruin, N.; Buell, M. Perceived or not perceived: Film character models for expressive nlg. In Proceedings of the International Conference on Interactive Digital Storytelling. Springer, 2011, pp. 109–121

  27. [35]

    Sunspring

    Goodwin, R.; Sharp, O. Sunspring. YouTube website. https://youtu. be/LY7x2Ihqjmc2016

  28. [36]

    Generating sentence planning variations for story telling

    Lukin, S.M.; Reed, L.I.; Walker, M.A. Generating sentence planning variations for story telling. arXiv preprint arXiv:1708.08580 2017

  29. [37]

    BLEU: a method for automatic evaluation of machine translation

    Papineni, K.; Roukos, S.; Ward, T.; Zhu, W.J. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 2002, pp. 311–318

  30. [38]

    Language models are unsupervised multitask learners

    Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. Language models are unsupervised multitask learners. OpenAI blog 2019, 1, 9

  31. [39]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, ...

  32. [40]

    Figures I; Vol

    Genette, G. Figures I; Vol. 1, Points, 1976

  33. [41]

    Narratology: Introduction to the theory of narrative; University of Toronto Press, 2009

    Bal, M.; Van Boheemen, C. Narratology: Introduction to the theory of narrative; University of Toronto Press, 2009

  34. [42]

    A statistical, grammar-based approach to microplanning

    Gardent, C.; Perez-Beltrachini, L. A statistical, grammar-based approach to microplanning. Computational Linguistics 2017, 43, 1–30

  35. [43]

    Summarization-inspired temporal-relation extraction: tense-pair templates and treebank-3 analysis 2007

    Dorr, B.; Gaasterland, T. Summarization-inspired temporal-relation extraction: tense-pair templates and treebank-3 analysis 2007

  36. [44]

    Towards a Mixed Evaluation Approach for Computational Narrative Systems

    Zhu, J. Towards a Mixed Evaluation Approach for Computational Narrative Systems. In Proceedings of the Proc. ICCC’12, 2012, pp. 150–154

  37. [45]

    Automated Evaluation of Meter and Rhyme in Russian Generative and Human-Authored Poetry 2025

    Koziev, I. Automated Evaluation of Meter and Rhyme in Russian Generative and Human-Authored Poetry 2025. [arXiv:cs.CL/2502.20931]

  38. [46]

    Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation, 2025, [arXiv:cs.CL/2502.13207]

    Franceschelli, G.; Musolesi, M. Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation, 2025, [arXiv:cs.CL/2502.13207]

  39. [47]

    Clases 1985: Algunos problemas de teoría literaria; Paidós, 2015

    Ludmer, J. Clases 1985: Algunos problemas de teoría literaria; Paidós, 2015

  40. [48]

    Austin JL-How to Do Things With Words.pdf, 1962

    Austin, J. Austin JL-How to Do Things With Words.pdf, 1962

  41. [49]

    La aproximación al texto literario en la enseñanza obligatoria

    Bertochi, D. La aproximación al texto literario en la enseñanza obligatoria. Textos de didáctica de la lengua y la literatura 1995, pp. 23–38

  42. [50]

    Educación y literatura; Mantaro, 2003

    Huamán, M.Á. Educación y literatura; Mantaro, 2003. 26 of 26

  43. [51]

    La telenovela mexicana: forma y contenido de un formato narrativo de ficcíon de alcance mayoritario

    Lizaur Guerra, M.B.d. La telenovela mexicana: forma y contenido de un formato narrativo de ficcíon de alcance mayoritario. PhD thesis, Universidad Nacional Autónoma de México. Facultad de Filosofía y Letras., 2003

  44. [52]

    Merchants of culture: the publishing business in the twenty-first century; John Wiley & Sons, 2013

    Thompson, J.B. Merchants of culture: the publishing business in the twenty-first century; John Wiley & Sons, 2013

  45. [53]

    La marca del editor; Anagrama, 2014

    Calasso, R. La marca del editor; Anagrama, 2014

  46. [54]

    The logic of narrative possibilities

    Bremond, C.; Cancalon, E.D. The logic of narrative possibilities. New Literary History 1980, 11, 387–411

  47. [55]

    Art as technique

    Shklovsky, V .; et al. Art as technique. Literary theory: An anthology 1917, 3

  48. [56]

    Attention is all you need

    Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in neural information processing systems, 2017, pp. 5998–6008

  49. [57]

    GPT2-spanish

    Oñate Latorre, A.; Ortiz Fuentes, J. GPT2-spanish. https://huggingface.co/DeepESP/gpt2-spanish, 2018

  50. [58]

    Language models are few-shot learners

    Brown, T.B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P .; Neelakantan, A.; Shyam, P .; Sastry, G.; Askell, A.; et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 2020

  51. [59]

    From novice to expert: Implications of language skills and writing-relevant knowledge for memory during the development of writing skill

    McCutchen, D. From novice to expert: Implications of language skills and writing-relevant knowledge for memory during the development of writing skill. Journal of writing research 2011, 3, 51–68

  52. [60]

    Sentence-bert: Sentence embeddings using siamese bert-networks

    Reimers, N.; Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084 2019

  53. [61]

    Semantic cosine similarity

    Rahutomo, F.; Kitasuka, T.; Aritsugi, M.; et al. Semantic cosine similarity. In Proceedings of the 7th international student conference on advanced science and technology ICAST. University of Seoul South Korea, 2012, Vol. 4, p. 1. Disclaimer/Publisher’s Note: The statements, o...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.