REVIEW 4 major objections 4 minor 21 references
Evaluating Quality of Gaming Narratives Co-created with AI
T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper claims that a validated list of 23 story quality dimensions, mapped onto Kano satisfaction categories, can guide game developers in deciding which quality aspects to secure when AI co-writes game narratives.
desk verdict A clearly written expert-driven checklist for AI narrative quality, but the Kano mapping is expert prediction, not player measurement, so the prioritization claim runs ahead of the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the list of 23 story quality dimensions (SQDs) taken from a systematic survey of story evaluation, together with two methodological instruments: the Delphi study and the Kano model. The Delphi study is an iterative, anonymous expert-consensus process; here it is used in its ranking-type form, with ten narrative-design experts scoring the importance of each SQD on a 1–5 Likert scale and assigning each SQD to a Kano category. The Kano model classifies product attributes by how their presence or absence influences satisfaction, yielding categories such as must-have, one-dimensional (performance), attractive (delighter), indifferent, and reverse. In this application, e
What would settle it
Run a standard Kano survey with functional/dysfunctional paired questions on the same 23 dimensions with a large sample of players for a specific game genre, and compare the resulting categories to the expert predictions; substantial divergence—for example, players rating several 'must-have' dimensions as indifferent—would refute the paper's central prioritization claim.
Extended reading notes
Core claim
The paper's central claim is that the initial list of 23 story quality dimensions is relevant and usable for prioritizing quality assurance in AI-co-created game narratives. Support comes from the first round of a ranking-type Delphi study: ten experts rated every dimension, none received a median importance below 3.0 on the 1–5 scale, 78% scored at least 3.5, and 26% scored above 4.5. Mapping the dimensions onto the Kano model, the experts predicted that 57% are one-dimensional (satisfaction proportional to how well the story performs the dimension), 26% are must-haves, 13% are attractive, and one is indifferent. The expert panel also flagged voice and genre alignment as quality dimensions
Load-bearing premise
The whole prioritization rests on ten experts' judgments about how players will feel—no actual player satisfaction data was collected, so if those judgments diverge from real player responses, the advice collapses.
Editorial extensions
If this is right
- Developers can treat the 23 SQD list as a validated baseline: every dimension survived expert review, so an AI-narrative quality rubric should cover all of them.
- With 57% of dimensions predicted to be one-dimensional, most quality aspects should be checked for gradient—higher performance is expected to yield proportionally higher satisfaction, and shortfalls should matter.
- The 26% classified as must-have are the non-negotiables: a story failing on these would dissatisfy players even if other dimensions are strong.
- Voice and genre alignment should be folded into future quality evaluations, as the experts identified them as missing from the literature-derived list.
- The categories give LLM-as-a-Judge prompts a concrete set of criteria: a judge model can be told which dimensions to score and which to weight as must-haves versus differentiators.
Reading between the lines
- The Kano assignment is expert prediction, not player measurement; a player-facing Kano survey on the same 23 dimensions could plausibly move items between categories, so the prioritization is provisional until such data exists.
- Several dimensions that are 'one-dimensional' by expert consensus, such as surprise or pacing, may actually have an inverted-U effect on player experience—too much can hurt—which would complicate the linear-satisfaction interpretation.
- Because the panel is small and geographically clustered, the predicted priorities may be specific to the expert culture sampled; testing against a large, genre-specific player survey would show how much the ordering generalizes.
- The framework lends itself to automation: each SQD can become a scoring criterion for an automated narrative judge, with the Kano category setting whether the criterion is a gating requirement or a bonus signal.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an evaluation framework for AI-generated game narratives. It compiles 23 story quality dimensions from an external literature survey, runs a first-round Delphi study with 10 narrative-design experts to rate the importance of each dimension, and asks the same experts to categorize each dimension into Kano satisfaction types (Delighter/One-dimensional/Must-be/Indifferent). The reported results are that all 23 dimensions are at least moderately important, 57% are classified as One-dimensional, 26% as Must-have, 13% as Attractive, and one as Indifferent. The authors conclude the list is relevant and can inform game developers' prioritization of quality aspects in LLM co-created narratives. A future second Delphi round and a large-scale player survey are mentioned in Section V.
Significance. If the central claim were fully supported, the paper would provide a practical, evidence-based checklist for developers assessing LLM-generated narrative quality, with the Kano mapping giving a satisfaction-based rationale for prioritization. The manuscript's strengths include grounding the initial 23 dimensions in a recent systematic survey of story evaluation, recruiting a geographically and gender-diverse expert panel, and explicitly planning a player-focused follow-up. However, the load-bearing step—the Kano classification—is derived from expert judgment rather than from players, and the Delphi evidence is limited to a single round. As presented, the paper is best read as a preliminary expert-consensus framework whose prioritization claims require validation; that validation is deferred to future work.
major comments (4)
- [Section III-C and Section IV] The Kano mapping is the central analytic step, but it is not executed per the Kano method. The paper states: "Rather than simply ranking these dimensions, our approach involves asking experts to classify each dimension within one of the Kano categories." Standard Kano classification requires paired functional/dysfunctional questions answered by end users (players), not direct expert categorization. Section IV then reports these categories as "expected" satisfaction effects, and Section V uses them to advise prioritization. Without player-derived validation, the Kano categories are unvalidated expert predictions, not measurements of player satisfaction. The authors should either add player data or explicitly and consistently reframe the Kano results as hypotheses, removing the prioritization advice from the central conclusion.
- [Section III-B and Section IV] The study is described as a Delphi study, but only one round is presented and no consensus metrics are reported. Delphi is an iterative process whose validity depends on convergence across rounds, typically assessed by dispersion measures (e.g., IQR), stability between rounds, and pre-specified stopping criteria. The paper reports medians and aggregate percentages but no per-dimension distributions, IQRs, or consensus indicators. With n=10, a median above 3.0 is weak evidence of relevance, especially when the paper claims that 78% and 26% of dimensions reached particular thresholds. The authors should report the full data (e.g., a table with medians, IQRs, and category counts) and either conduct additional rounds or label the work as a first-round expert consultation rather than a completed Delphi.
- [Section V-A] The two emergent dimensions, voice and genre alignment, are introduced based on comments from the same expert panel and are then treated as additions to the framework, with genre alignment classified as a "must-be" requirement. Because these dimensions were not part of the originally validated list, they have not undergone the same Delphi rating or Kano categorization process. The statements about their importance and Kano category are expert opinions from at most a subset of the panel, not results of the study's methodology. This should be clearly labeled as hypothesis generation, not validated output.
- [Section IV, Figure 2] Figure 2, which is supposed to contain the per-dimension importance scores and Kano categories, is not included in the manuscript text; only its caption appears. Consequently, the reader cannot verify the reported percentages (57%, 26%, 13%) or examine the classification of individual dimensions. A table or figure showing all 23 dimensions, their median importance, distribution, and Kano category is necessary to support the claims.
minor comments (4)
- [Section IV] Terminology is inconsistent: the paper alternates between "Must-have" and "Most-haves," and between "Delighter" and "Attractive." Use the standard Kano terms consistently.
- [Section V] Typo in "second round of this Dephi study"—should be "Delphi." Also, the abbreviation "SQDs" is sometimes used without definition in figure captions and text.
- [References] References [19] and [20] have incomplete venue information and appear to contain placeholder years ("1984"). Please correct.
- [Section V] The sentence "In a forthcoming paper, we conduct a second round of this Dephi study" would be better phrased as "future work" and the manuscript's status as a preliminary report should be stated in the abstract or introduction.
Circularity Check
No significant circularity: the 23 dimensions come from an external literature survey, the Delphi ratings are new expert data, and the Kano categories are explicitly labeled as expert predictions rather than player measurements.
full rationale
The paper's derivation chain is linear and self-contained. The initial 23 SQDs are taken from an external systematic survey (Yang & Jin [15]), not from the present authors' prior work. The Delphi study collects new expert importance ratings and expert Kano classifications (Section III-B/C). The results report those ratings and classifications (Section IV), and the conclusion (Section V) infers relevance and prioritization value from them. No step defines its output in terms of its own output: the Kano categories are not derived from player-satisfaction data, nor is any fitted parameter relabeled as a prediction. The paper explicitly calls the Kano assignments 'predicted Kano category of player satisfaction' and states that a forthcoming large-scale survey will capture players' preferences, acknowledging that player data are not yet collected. That is a limitation on external validity/generalizability, not a circular derivation. There are no load-bearing self-citations; references [15] and [18] are external sources. Therefore the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The 23 SQDs from Yang and Jin's survey are a comprehensive starting list for story quality in AI narratives.
- ad hoc to paper Expert categorization into Kano categories predicts player satisfaction.
- ad hoc to paper Kano categories can be assigned by direct expert categorization without functional/dysfunctional paired questions.
- domain assumption A single Delphi round with n=10 experts yields stable conclusions.
Cite this review
Pith. "Pith review of Evaluating Quality of Gaming Narratives Co-created with AI." pith.science (2026). https://pith.science/paper/2HPNZ7JG
@misc{pith2026250904239,
author = {Pith},
title = {Pith review of: Evaluating Quality of Gaming Narratives Co-created with AI},
year = {2026},
howpublished = {\url{https://pith.science/paper/2HPNZ7JG}},
note = {Machine review of arXiv:2509.04239}
}
read the original abstract
This paper proposes a structured methodology to evaluate AI-generated game narratives, leveraging the Delphi study structure with a panel of narrative design experts. Our approach synthesizes story quality dimensions from literature and expert insights, mapping them into the Kano model framework to understand their impact on player satisfaction. The results can inform game developers on prioritizing quality aspects when co-creating game narratives with generative AI.
Figures
Reference graph
Works this paper leans on
-
[1]
M. J. Nelson, N. Shaker, and J. Togelius, Procedural Content Generation in Games. Springer Cham, 2016
work page 2016
-
[2]
“Yoli games,” YOLI ApS, May 2025, accessed: 2025-05-24. [Online]. Available: https://www.playyoli.com
work page 2025
-
[3]
L. M. Csepregi, “The effect of context-aware llm-based npc conver- sations on player engagement in role-playing video games,” 2021, unpublished manuscript
work page 2021
-
[4]
A framework for exploring player perceptions of llm-generated dialogue in commercial video games,
N. Akoury, Q. Yang, and M. Iyyer, “A framework for exploring player perceptions of llm-generated dialogue in commercial video games,” in Findings of the Association for Computational Linguistics: EMNLP 2023, 2023, pp. 2295–2311
work page 2023
-
[5]
Generating video game scripts with style,
G. L. Latouche, L. Marcotte, and B. Swanson, “Generating video game scripts with style,” in Proceedings of the 5th Workshop on NLP for Conversational AI (NLP4ConvAI 2023) , 2023, pp. 129–139
work page 2023
-
[6]
Chatter generation through language models,
M. M ¨uller-Brockhausen, G. Barbero, and M. Preuss, “Chatter generation through language models,” in 2023 IEEE Conference on Games (CoG) . IEEE, 2023, pp. 1–6
work page 2023
-
[7]
Fictional worlds, real connections: Developing community storytelling social chatbots through llms,
Y . Sun, H. Wang, P. M. Chan, M. Tabibi, Y . Zhang, H. Lu, Y . Chen, C. H. Lee, and A. Asadipour, “Fictional worlds, real connections: Developing community storytelling social chatbots through llms,” 2023
work page 2023
-
[8]
T. Ashby, B. K. Webb, G. Knapp, J. Searle, and N. Fulda, “Personalized quest and dialogue generation in role-playing games: A knowledge graph- and language model-based approach,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. ACM, 2023, article 290
work page 2023
Show all 21 references
-
[9]
Towards grounded dialogue gener- ation in video game environments,
N. Akoury, R. Salz, and M. Iyyer, “Towards grounded dialogue gener- ation in video game environments,” 2023
2023
-
[10]
Language as reality: A co-creative storytelling game experience in 1001 nights using generative ai,
Y . Sun, Z. Li, K. Fang, C. H. Lee, and A. Asadipour, “Language as reality: A co-creative storytelling game experience in 1001 nights using generative ai,” in Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment , vol. 19, 2023, p...
2023
-
[11]
From playing the story to gaming the system: Repeat experiences of a large language model-based interactive story,
Q. R. Yong and A. Mitchell, “From playing the story to gaming the system: Repeat experiences of a large language model-based interactive story,” in International Conference on Interactive Digital Storytelling . Springer, 2023, pp. 395–409
2023
-
[12]
The chronicles of chatgpt: Generating and evaluating visual novel narratives on climate change through chatgpt,
M. C. Gursesli, P. Taveekitworachai, F. Abdullah, M. F. Dewantoro, A. Lanata, A. Guazzini, V . K. L ˆe, A. Villars, and R. Thawonmas, “The chronicles of chatgpt: Generating and evaluating visual novel narratives on climate change through chatgpt,” in International Conference o...
2023
-
[13]
What is waiting for us at the end? inherent biases of game story endings in large language models,
P. Taveekitworachai, F. Abdullah, M. C. Gursesli, M. F. Dewantoro, S. Chen, A. Lanata, A. Guazzini, and R. Thawonmas, “What is waiting for us at the end? inherent biases of game story endings in large language models,” in International Conference on Interactive Digital Storyte...
2023
-
[14]
Journey of chatgpt from prompts to stories in games: the positive, the negative, and the neutral,
P. Taveekitworachai, M. C. Gursesli, F. Abdullah, S. Chen, F. Cala, A. Guazzini, A. Lanata, and R. Thawonmas, “Journey of chatgpt from prompts to stories in games: the positive, the negative, and the neutral,” in 2023 IEEE 13th International Conference on Consumer Electronics-...
2023
-
[15]
What makes a good story and how can we measure it? a comprehensive survey of story evaluation,
D. Yang and Q. Jin, “What makes a good story and how can we measure it? a comprehensive survey of story evaluation,” arXiv preprint arXiv:2408.14622, 2024
2024 arXiv
-
[16]
A survey on llm-as-a-judge,
J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y . Shen, S. Ma, H. Liu, S. Wang, K. Zhang, Y . Wang, W. Gao, L. Ni, and J. Guo, “A survey on llm-as-a-judge,” 2025. [Online]. Available: https://arxiv.org/abs/2411.15594
2025 arXiv
-
[17]
The delphi method and its contribution fo decision-making
E. Ziglio, “The delphi method and its contribution fo decision-making.” in Gazing Into the Oracle: The Delphi Method and Its Application to Social Policy and Public Health , M. Adler and E. Ziglio, Eds. Jessica Kingsley Publishers, 1996, pp. 3–26
1996
-
[18]
Attractive quality and must-be quality,
N. Kano, N. Seraku, F. Takahashi, and S. ichi Tsuji, “Attractive quality and must-be quality,” Journal of The Japanese Society for Quality Control, vol. 14, no. 2, pp. 147–156, 1984
1984
-
[19]
Bayesian modelling of the well-made surprise,
P. Chieppe, P. Sweetser, and E. Newman, “Bayesian modelling of the well-made surprise,” ?, 1984
1984
-
[20]
Predicting grammaticality on an ordinal scale,
M. Heilman, A. Cahill, N. Madnani, M. Lopez, M. Mulholland, and J. Tetreault, “Predicting grammaticality on an ordinal scale,” ?, 1984
1984
-
[21]
H. A. Linstone and M. Turoff, The Delphi Method: Techniques and Applications, 2nd ed. Reading, MA: Addison-Wesley, 2002
2002
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.