Pith. sign in

REVIEW 3 major objections 7 minor 38 references

The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing

T0 review · 3 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Conflicting AI-vs-human literary scores are explained by two stable reader profiles, the paper argues: surface-focused lay readers and holistic experts.

desk verdict A transparent, methodologically novel perspective-aware evaluation study whose two-profile finding is undermined by a reader-type/corpus confound; worth reviewing for the framework, not yet for the result. read the letter →

arxiv 2506.03310 v1 pith:YBUZXY2Z submitted 2025-06-03 cs.CL cs.HC

classification cs.CLcs.HC
keywords AIcreativewritingliteraryevaluationreaderpreferencestextualfeaturesRandomForestfeatureimportanceperspectivistexpertvsnon-expertreadersstorygeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the conflicting results in studies of AI-generated versus human-authored stories are not measurement noise or evidence of intrinsic text quality; they reflect genuine differences in what readers value. The authors model each reader's judgments with a separate classifier over 17 textual features, producing a per-reader feature importance vector, and find that these vectors cluster into two profiles: surface-focused readers, mainly non-experts, who weight readability and linguistic richness, and holistic readers, mainly literary experts, who weight thematic development, rhetorical variety, and sentiment dynamics. The conclusion is that whether AI writing beats human writing depends on which reader profile does the judging, so literary quality is relational rather than an objective property of the text.

What carries the argument

The central object is the per-reader feature importance vector, derived as the mean decrease in impurity (MDI) from an individually trained Random Forest classifier fed with feature differences between pairs of texts and the reader's binary preference. This vector is the empirical proxy for the weights in an assumed additive utility model $U_j(x_i)=\sum_{f} w_{jf} x_{if}$. k-means clustering of these vectors in a shared preference space exposes the two reader profiles. The vocabulary is a fixed set of 17 reference-less textual features covering coherence, readability, lexical diversity, sentiment dynamics, originality approximated by language-model log-likelihood, and thematic structure.

What would settle it

Give one set of stories (for example, the SLM synopses or the TTCW quartets) to a mixed panel of literary experts and lay readers under blind conditions, rerun the per-reader Random Forest and clustering pipeline, and look at whether importance vectors still separate by expertise. If they separate on the same corpus, the profile split is a reader effect; if they do not, the clusters are an artifact of which texts each group happened to rate.

Watch

Extended reading notes

Core claim

The central claim is that measurements of literary quality are a function of how text features align with each reader's preferences. Concretely, per-reader feature importance vectors, estimated from Random Forests trained on each reader's pairwise preferences, cluster into two groups that align with reader background: non-experts emphasize readability, syntactic depth, sentence length, and lexical diversity, while experts emphasize thematic entropy, rhetorical variety, sentence rhythm, and sentiment dynamics. Across five corpora, AI preference rates track this split: in lay-reader datasets where AI texts fall in the feature region those readers reward, AI preference is high, and in expert datasets where human texts occupy the valued region, AI preference is low. The paper therefore attributes the contradictory win rates in earlier work to the evaluative profile of the readers, not to an intrinsic quality gap.

Load-bearing premise

The load-bearing premise is that per-reader feature importance vectors measure stable reader preferences, but the datasets conflate reader background with the texts they rated: experts only saw high-literary human comparisons and non-experts only saw average human comparisons, so the two clusters could be corpus artifacts rather than reader types.

Editorial extensions

If this is right

  • Aggregate AI-vs-human win rates from different studies can disagree without any study being wrong; the disagreement is encoded in the reader mix behind each number.
  • A text's feature profile determines its reception: AI stories win when they land in the feature space their judges reward, and lose when they do not.
  • Per-reader preference models trained on pairwise judgments can be compared across datasets in a shared preference space, allowing reader profiles to be studied without asking readers to self-report their expertise.
  • Expert and non-expert judgments are not noisy versions of one true quality score; they are two different utility functions over textual features.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable next step the paper does not run: have literary experts and lay readers rate the same stories and check whether their importance vectors still separate, separating reader preferences from corpus artifacts.
  • If the account holds, evaluation reports should state the reader mix; otherwise contradictory AI-vs-human headlines will keep appearing as artifacts of who judged.
  • The same per-reader importance pipeline could be applied to other contested quality judgments, where disagreements may decompose into stable evaluator profiles rather than noise.
  • A causal extension would generate stories optimized for each profile and check whether preferences flip, turning the observed correlation into a manipulable variable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper proposes a perspectivist explanation for conflicting findings on AI versus human creative writing. It uses five public datasets (1,471 stories, 101 annotators), extracts 17 reference-less textual features, trains per-reader Random Forest classifiers on pairwise preference differences, and uses MDI feature importances as proxies for reader utility weights. The authors cluster readers into two profiles—a 'surface-focused' profile mainly composed of lay readers and a 'holistic' profile mainly composed of experts—and argue that the conflicting AI-preference rates across prior studies are explained by these different reader profiles rather than by intrinsic text quality. The paper is clearly written, ships code and data, and is framed as a contribution to reader-sensitive evaluation of text generation.

Significance. If the central identification were secure, this would be a substantial contribution: it offers an interpretable and reproducible framework for reconciling contradictory evaluation results in AI creative writing, and it extends the perspectivist research program to individual-level preference modeling. Strengths include the open-source release of code, data, and results; the use of five public datasets; the careful selection of 17 low-redundancy textual features; and the fact that clustering is performed without reader metadata, so the expertise correlation is emergent. The main weakness is that reader expertise is nested within datasets: experts and lay readers evaluate different corpora with different AI/human text contrasts. This makes the central claim—that the two emergent clusters are stable reader profiles—currently under-supported. The paper's own limitation section acknowledges observational data and approximate expertise categories, but it does not address the structural corpus-reader confound.

major comments (3)
  1. [§4.6, §5.2, Fig. 2] Reader type is nested within dataset: experts (TTCW, PRONVSPROMPT) and lay readers (SLM, HANNA) never evaluate the same texts, and CONFEDERACY provides only literature students. The MDI importances in §4.6 summarize which features best separate the AI from the human texts in the corpus each reader saw; if a corpus happens to differ strongly on log-likelihood, sentiment variance, or rhetorical variety, those features will have high MDI regardless of a reader's stable preferences. The cluster separation in Fig. 2 therefore coincides with dataset membership and cannot, as it stands, support the conclusion that two reader profiles exist. Section 7's observational-data caveat does not address this structural confound. A concrete remedy is to compare readers of different expertise on the same corpus, or to cluster readers separately within each dataset and show that the same two profiles emerge, or to compare the observed clusters with a null distribution obtained by permuting corpus labels while holding reader labels fixed.
  2. [§4.5, Appendix C, Table 2] Per-reader training sets are very small. For example, CONFEDERACY has 10 raters and 65 texts, with each rater reviewing 13 texts, so at most 78 unique pairs (156 instances after balancing) are available per reader; for TTCW, the text says each story was evaluated by three experts, but Table 2 reports only 7 readers while Appendix C.4 names 10 experts. With 17 features, MDI estimates from Random Forests on such small samples will be high-variance. The paper reports F1 scores in Appendix B but does not report uncertainty for MDI or show that the two clusters are stable under resampling of readers or pairs. Without such evidence, the two clusters may reflect estimation noise rather than genuine reader types. Please report per-reader MDI confidence intervals (e.g., via bootstrap) or use a lower-capacity model, and verify cluster stability.
  3. [§5.3 and Eq. (3)] MDI is unsigned and reflects discriminative power in the pairwise classification task, not the sign or direction of a reader's preference. The interpretation of Cluster 0 as valuing 'elaborate and lexically rich prose presented within an accessible framework' and Cluster 1 as valuing 'thematic development, expressive stylistics, and sentiment dynamics' is obtained by cross-referencing unsigned MDI values with mean feature values in Table 3 for texts preferred by each cluster. Because those preferred texts come from different corpora with different AI/human contrasts, the directional reading is again confounded by corpus. A cleaner approach would be to compute directional preference weights directly, for example from a linear model on pairwise differences, or by comparing the feature values of preferred versus non-preferred texts within each reader's own evaluation set, and only then clustering the directional weights.
minor comments (7)
  1. [§5.2] The text states that the number of clusters was 'determined using standard validation techniques (see Section 4.6)', but Section 4.6 does not describe any cluster validation; please add the actual criterion (e.g., silhouette or elbow) and the resulting values.
  2. [Table 2 vs. Appendix C.4] Table 2 gives 7 readers for TTCW, while Appendix C.4 describes 10 experts; please reconcile this count and report the exact number of pairwise preference instances per reader for each dataset.
  3. [Eq. (2)] Equation (2) introduces an epsilon term without defining it; please state explicitly that it is a small constant added for numerical stability, or remove it.
  4. [Appendix A] Equation (14) in Appendix A is an empty line; if this is a formatting artifact, remove it and renumber the subsequent equations.
  5. [Table 3] The column header 'Max subord...' is truncated; use the full feature name 'Max subordination depth' for all tables.
  6. [§4.1] In the SLM corpus description, the phrase 'regarding the author's' appears incomplete; it should read 'regarding the author's identity' or similar.
  7. [Fig. 2b] The stacked bars in Figure 2b would benefit from an explicit legend defining the cluster colors; the numeric labels are helpful but the colors are not identified in the caption.

Circularity Check

1 steps flagged · score 2.0 of 10

No substantial circularity; one definitional explanatory use of reader-preference centroids in RQ1 keeps the score at 2.

  1. self definitional [Section 5.1 (SLM discussion), Eq. (2)]
    "x∗ j = P xi∈X top j ρj(xi) xi / P xi∈X top j ρj(xi) + ϵ ... Given that reader preferences are densely clustered within this same overlapping area, the AI's greater presence in this zone—where texts may be perceived as indistinguishable or even preferable according to the metrics—likely explains its higher preference rate of 57.59%."

    The reader preference centroid in Eq. (2) is defined as the preference-weighted average of features from each reader's own top-rated texts. Its position in feature space therefore restates which texts that reader scored highest. Section 5.1 then uses centroid alignment with the AI zone to explain the reader's AI-preference rate, which is circular by construction: the centroid already encodes the preference scores that produce the rate. The explanation adds no independent evidence beyond re-describing the input ratings. This step is local to RQ1 and is not the load-bearing derivation for the two-cluster reader profile result in Section 5.2.

full rationale

The paper's central derivation chain is empirical rather than derivational: per-reader Random Forests are trained on pairwise preference comparisons, MDI feature importances are extracted, and these importance vectors are clustered without using reader metadata. The resulting clusters are emergent and are not constrained by construction to match expertise labels, so the two-profile claim has independent empirical content. The two self-cited datasets (SLM and PRONVSPROMPT) are used as published data sources, not as authority for the paper's assumptions, so this is normal self-citation rather than load-bearing circularity. No uniqueness theorem, ansatz-smuggling citation, or fitted-parameter-then-prediction chain is present. The one definitional step is the use of reader preference centroids in Section 5.1 to 'explain' AI-preference rates: the centroids are literally constructed from the same preference scores, making that explanation a restatement of the input. The broader concern that expert and lay readers are confounded with different corpora is a validity threat, not a circularity-by-construction finding, and does not raise the circularity score under the stated criteria. Overall, the central claim is self-contained against external benchmarks, with one minor self-definitional interpretation keeping the score at 2.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the assumptions that feature importances from per-reader Random Forests reflect stable reader preferences, that the 17-feature set is adequate, and that the five datasets can be pooled; the main threat is the dataset-reader confound.

free parameters (4)
  • k (number of reader clusters) = 2
    k-means with k=2 chosen via 'standard validation techniques' (Section 4.6), not detailed; the two-cluster structure is the paper's central descriptive result.
  • Top-25% threshold for reader preference centroids = 0.25
    Eq. 2 defines the reader preference centroid over the top 25% of a reader's evaluations; this cutoff is stated without sensitivity analysis.
  • LDA topic count K* = argmax C_v over [Kmin, Kmax]
    Eq. 4 selects the number of LDA topics per corpus by maximizing topic coherence; this data-dependent choice affects topic entropy and density features.
  • Random Forest hyperparameters = grid-searched per reader
    n_estimators, max_depth, min_samples_split, min_samples_leaf are tuned by grid search per reader (Appendix B); per-reader tuning on tiny samples adds variance to importances.
assumptions (5)
  • domain assumption Each reader's evaluation can be approximated as an additive utility function over 17 intrinsic textual features (Eq. 3).
    Stated in Section 3 as a conceptual starting point; the paper itself notes practical estimation uses non-linear Random Forest proxies, but the interpretability of feature importances as preferences depends on this approximation.
  • domain assumption MDI feature importances from a per-reader Random Forest approximate the conceptual weight vector w_j.
    Section 4.6 states these serve as empirical proxies; MDI is known to be biased when features are correlated, and the paper only checks pairwise correlations below 0.7, not higher-order interactions.
  • domain assumption Min-max normalized per-reader scores rho_j produce pairwise preference labels that reflect genuine preferences rather than scale artifacts.
    Section 4.3 aggregates heterogeneous Likert and binary judgments into a single score per reader, then discards ties; this assumes comparability across dimensions within a reader.
  • domain assumption The five datasets can be compared in a shared preference space defined by feature importances, despite differences in prompts, genres, AI models, and annotation protocols.
    Section 5.4 claims cross-dataset consistency; the pooling assumes that MDI vectors are commensurable across corpora, which is not tested explicitly.
  • domain assumption Length-unadjusted raw log-likelihood under Llama-3.2-3B is a valid reference-less originality proxy.
    Appendix A.2 defines OrigLL as a sum over tokens without per-token normalization; longer texts will accumulate more negative terms, confounding originality with length.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing." pith.science (2026). https://pith.science/paper/YBUZXY2Z

@misc{pith2026250603310,
  author       = {Pith},
  title        = {Pith review of: The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YBUZXY2Z}},
  note         = {Machine review of arXiv:2506.03310}
}
read the original abstract

Recent studies comparing AI-generated and human-authored literary texts have produced conflicting results: some suggest AI already surpasses human quality, while others argue it still falls short. We start from the hypothesis that such divergences can be largely explained by genuine differences in how readers interpret and value literature, rather than by an intrinsic quality of the texts evaluated. Using five public datasets (1,471 stories, 101 annotators including critics, students, and lay readers), we (i) extract 17 reference-less textual features (e.g., coherence, emotional variance, average sentence length...); (ii) model individual reader preferences, deriving feature importance vectors that reflect their textual priorities; and (iii) analyze these vectors in a shared "preference space". Reader vectors cluster into two profiles: 'surface-focused readers' (mainly non-experts), who prioritize readability and textual richness; and 'holistic readers' (mainly experts), who value thematic development, rhetorical variety, and sentiment dynamics. Our results quantitatively explain how measurements of literary quality are a function of how text features align with each reader's preferences. These findings advocate for reader-sensitive evaluation frameworks in the field of creative text generation.

Figures

Figures reproduced from arXiv: 2506.03310 by the authors.

Figure 1
Figure 1. PCA projection of textual feature vectors for AI-generated and human-authored stories across four datasets [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Clustering of users based on feature impor [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Radar chart illustrating the mean Random [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Comparison of F1 scores on test set for Random Forest and Logistic Regression. The boxplots illustrate [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]
Figure 5
Figure 5. Figure 5: Correlation matrix of the textual features developed across all corpora. Notably, no highly positive or [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

38 extracted references · 18 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    AI at Meta . 2024. https://arxiv.org/abs/2404.11225 The Llama 3 Herd of Models . Preprint, arXiv:2404.11225

  4. [4]

    Regina Barzilay and Mirella Lapata. 2008. https://doi.org/10.1162/coli.2008.34.1.1 Modeling local coherence: An entity-based approach . Computational Linguistics, 34(1):1--34

  5. [5]

    Yuri Bizzoni, Pascale Feldkamp, Kristoffer Nielbo, Ida-Marie Lassen, and Mads Thomsen. 2024 a . https://repository.clarin.dk/repository/xmlui/handle/20.500.12115/56 FabulaNet Literary Quality Dataset . https://centre-for-humanities-computing.github.io/fabula-net/. Accepted: 2024-09-02T08:53:01Z Publisher: University Århus

  6. [6]

    Yuri Bizzoni, Pascale Feldkamp, and Kristoffer Laigaard Nielbo. 2024 b . http://www.scopus.com/inward/record.url?scp=85210802067&partnerID=8YFLogxK Global Coherence , Local Uncertainty : 2024 Computational Humanities Research Conference , CHR 2024 . Proceedings of the Computational Humanities Research Conference 2024, 3834:172--185. Publisher: CEUR-WS

  7. [7]

    Yuri Bizzoni, Ida Marie Lassen, Telma Peura, Mads Rosendahl Thomsen, and Kristoffer Nielbo. 2022 a . https://aclanthology.org/2022.nlperspectives-1.3/ Predicting Literary Quality How Perspectivist Should We Be ? In Proceedings of the 1st Workshop on Perspectivist Approaches to NLP @ LREC2022 , pages 20--25, Marseille, France. European Language Resources A...

  8. [8]

    Yuri Bizzoni, Kristoffer Laigaard Nielbo, and Mads Rosendahl Thomsen. 2022 b . https://doi.org/10.18653/v1/2022.nlp4dh-1.5 Fractality of sentiment arcs for literary quality assessment: The case of Nobel laureates . In Proceedings of the 2nd International Workshop on Natural Language Processing for Digital Humanities , pages 31--41, Taipei, Taiwan. Associa...

Show all 38 references
  1. [9]

    Blei, Andrew Y

    David M. Blei, Andrew Y. Ng, and Michael I. Jordan. 2003. Latent dirichlet allocation. Journal of Machine Learning Research, 3:993--1022

  2. [10]

    Joanne Boisson, Zara Siddique, Hsuvas Borkakoty, Dimosthenis Antypas, Luis Espinosa Anke, and Jose Camacho-Collados. 2025. https://aclanthology.org/2025.coling-main.448/ Automatic extraction of metaphoric analogies from literary texts: Task formulation, dataset construction, a...

  3. [11]

    Leo Breiman. 2001. Random forests. Machine Learning, 45(1):5--32

  4. [12]

    Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, and et al

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, and et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33

  5. [13]

    Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020. Evaluation of text generation: A survey. arXiv preprint arXiv:2006.14799

  6. [14]

    Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu. 2024. https://doi.org/10.1145/3613904.3642731 Art or Artifice ? Large Language Models and the False Promise of Creativity . In Proceedings of the 2024 CHI Conference on Human Factors in ...

  7. [15]

    Suchanek, and Chloé Clavel

    Cyril Chhun, Pierre Colombo, Fabian M. Suchanek, and Chloé Clavel. 2022. https://aclanthology.org/2022.coling-1.509/ Of Human Criteria and Automatic Metrics : A Benchmark of the Evaluation of Story Generation . In Proceedings of the 29th International Conference on Computation...

  8. [16]

    Suchanek, and Chloé Clavel

    Cyril Chhun, Fabian M. Suchanek, and Chloé Clavel. 2024. https://doi.org/10.1162/tacl_a_00689 Do Language Models Enjoy Their Own Stories ? Prompting Large Language Models for Automatic Story Evaluation . Transactions of the Association for Computational Linguistics, 12:1122--1142

  9. [17]

    Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. GoEmotions : A dataset of fine-grained emotions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4040--4054. Association f...

  10. [18]

    Pascale Feldkamp, Yuri Bizzoni, Mads Rosendahl Thomsen, and Kristoffer L. Nielbo. 2024. https://doi.org/10.48694/jcls.3908 Measuring Literary Quality . Proxies and Perspectives . Journal of Computational Literary Studies, 3(1). Number: 1 Publisher: Universitäts- und Landesbibl...

  11. [19]

    Carlos Gómez-Rodríguez and Paul Williams. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.966 A Confederacy of Models : a Comprehensive Evaluation of LLMs on Creative Writing . In Findings of the Association for Computational Linguistics : EMNLP 2023 , pages 14504--14528...

  12. [20]

    Jochen Hartmann, Mark Heitmann, Christian Siebert, and Christina Schamp. 2023. More than a feeling: Accuracy and application of sentiment analysis. International Journal of Research in Marketing, 40(1):75--87

  13. [21]

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. https://doi.org/10.18653/v1/2020.acl-main.703 BART : Denoising sequence-to-sequence pre-training for natural language generation, translatio...

  14. [22]

    Jiwei Li and Eduard Hovy. 2014. A model of coherence based on distributed sentence representation. In Proc.\ EMNLP, pages 2039--2048

  15. [23]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa : A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692

  16. [24]

    Xiaofei Lu. 2010. https://doi.org/10.1075/ijcl.15.4.02lu Automatic analysis of syntactic complexity in second language writing . International Journal of Corpus Linguistics, 15(4):474--496

  17. [25]

    Ximing Lu, Zhaoyang Yu, Yuchen Sun, Yichen Zhang, and Nelson F. Liu. 2024. https://arxiv.org/abs/2410.04265 AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text . arXiv preprint arXiv:2410.04265

  18. [26]

    Guillermo Marco, Julio Gonzalo, M.Teresa Mateo-Girona, and Ramón Del Castillo Santos. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.1096 Pron vs Prompt : Can Large Language Models already Challenge a World - Class Fiction Author at Creative Text Writing ? In Proceedings of...

  19. [27]

    Guillermo Marco, Luz Rello, and Julio Gonzalo. 2025. https://aclanthology.org/2025.coling-main.437/ Small Language Models can Outperform Humans in Short Creative Writing : A Study Comparing SLMs with Humans and LLMs . In Proceedings of the 31st International Conference on Comp...

  20. [28]

    Philip M McCarthy and Scott Jarvis. 2010. Mtld, vocd-d, and hd-d: A validation study of sophisticated approaches to lexical diversity assessment. Behavior research methods, 42(2):381--392

  21. [29]

    McLaughlin

    Harry G. McLaughlin. 1969. Smog grading—a new readability formula. Journal of Reading, 12(8):639--646

  22. [30]

    Pascale Moreira, Yuri Bizzoni, Kristoffer Nielbo, Ida Marie Lassen, and Mads Thomsen. 2023. https://doi.org/10.18653/v1/2023.wnu-1.5 Modeling Readers ' Appreciation of Literary Narratives Through Sentiment Arcs and Semantic Profiles . In Proceedings of the 5th Workshop on Narr...

  23. [31]

    Mark E. J. Newman. 2010. Networks: An Introduction. Oxford University Press

  24. [32]

    Fabian Pedregosa, Ga \"e l Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. 2011. Scikit-learn: Machine learning in python. the Journal of machine Learning research, 12:2825--2830

  25. [33]

    Problem

    Barbara Plank. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.731 The “ Problem ” of Human Label Variation : On Ground Truth in Data , Modeling and Evaluation . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages 10671--10682, A...

  26. [34]

    Brian Porter and Edouard Machery. 2024. https://doi.org/10.1038/s41598-024-76900-1 AI -generated poetry is indistinguishable from human-written poetry and is rated more favorably . Scientific Reports, 14(1):26133. Publisher: Nature Publishing Group

  27. [35]

    Michael R \"o der, Andreas Both, and Alexander Hinneburg. 2015. Exploring the space of topic coherence measures. In Proc.\ WSDM, pages 399--408

  28. [36]

    Claude E. Shannon. 1948. A mathematical theory of communication. Bell System Technical Journal, 27:379--423

  29. [37]

    Evangelos Triantaphyllou. 2000. Multi-criteria decision making methods. Springer

  30. [38]

    Tianshu Yu, Ting-En Lin, Yuchuan Wu, Min Yang, Fei Huang, and Yongbin Li. 2025. https://doi.org/10.1162/tacl_a_00746 Diverse AI Feedback For Large Language Model Alignment . Transactions of the Association for Computational Linguistics, 13:392--407

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.