{"id":"f8d38b34-c1b4-444f-9f42-3f2b4fbf31f3","arxiv_id":"2508.02132","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A procedural game generation system that uses Rise/Fall emotional arcs to shape LLM-written branching stories and entity difficulty showed higher player enjoyment in a small ARPG user study.","lead":"This paper builds a game generation pipeline where classic emotional story arcs (Rise, Fall) guide how an LLM writes branching story nodes and sets enemy and item stats. In a 16-player ARPG test, arc-guided stories scored higher on enjoyment, but the evidence is thin and the paper overstates its results.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Difficulty control is internally contradictory: Method and Case Study say arcs adjust entity stats, Evaluation says stats were held constant, and Limitations concedes the confound; the only significant Enjoyment effect may not be attributable to emotional arcs.","rationale":"The reader's weakest assumption concerns the sentiment analysis showing that monotonic Rise and Fall arcs do not produce reliable directional emotional trajectories. That is a real and well-supported concern, and it does undercut the claim that all six arcs are faithfully realized. However, it primarily affects RQ2 and the 'emotional impact' portion of the claim; it does not directly invalidate the Enjoyment comparison. The difficulty-control contradiction is more load-bearing because it threatens the causal attribution of the only statistically significant result supporting the strongest claim. If enemy difficulty varied with the arc condition, the reported Enjoyment improvement cannot be assigned to emotional arc labels. The paper's own text makes this threat concrete: the system design says difficulty is derived from the arc, the evaluation says difficulty was held constant, and the limitations section admits the baseline lacked consistent difficulty modulation and that ranged enemies were present in both conditions. This internal inconsistency is exactly the kind of hidden confound that a concrete entity-statistics check can settle. I still do not think the paper should be rejected outright: the system is built, the study is transparent, and the qualitative findings are useful. But the central causal sentence in the abstract needs either corrected data or a substantially more guarded formulation, so the reader's CONDITIONAL verdict remains appropriate. My recommendation is therefore UNCHANGED relative to the reader's verdict, with the added requirement that the difficulty-control contradiction be resolved before the causal claim is accepted.","tokens_in":10595,"tokens_out":5120,"duration_ms":64729,"concrete_test":"Extract the generated JSON for all Baseline and Emotional Arc episodes used in the user study (rerunning the pipeline with the same user prompts and seeds if the logs are unavailable), then compare enemy attack, health, and hasRangedAttack distributions by condition and by node progression. If Fall/Arc nodes show systematically higher or more variable entity stats than Baseline nodes, the 'difficulty held constant' statement is false and the Enjoyment difference is confounded. A matched-stats rerun, or an analysis of Enjoyment with difficulty ratings or enemy stats as a covariate, would be required to determine whether emotional arc labels alone carry the effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal claim—that emotional arc integration significantly enhances engagement—rests on the user study's comparison between Baseline and Emotional Arc episodes. That comparison is undermined by an internal contradiction about the independent variable. The Method states that entity statistics are automatically configured based on the emotional arc, and the Case Study reports that enemies in Fall nodes were more likely to have higher attack, health, and ranged capabilities. The Evaluation, however, claims that 'gameplay mechanics, and difficulty parameters were held constant across conditions to isolate the narrative structure as the independent variable.' The Limitations then concedes that 'the Baseline condition did not include consistent difficulty modulation, and the presence of ranged-attack enemies in both conditions may have confounded comparative analysis of difficulty.' If enemy stats differed systematically between conditions, the significant Enjoyment result (p = 0.016, r = 0.789) could reflect difficulty pacing or enemy composition rather than the emotional arc labels themselves. The paper does not report per-condition entity statistics, so the confound cannot be resolved from the text. This is not a matter of overclaiming alone; it is an unresolved threat to the causal interpretation of the only statistically significant quantitative result.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an LLM-based pipeline for procedural game narrative generation in which branching story graphs are labeled with Rise/Fall emotional arcs drawn from Reagan et al. (2016). Each story node is automatically populated with narrative text, characters, items, and gameplay attributes, with enemy difficulty purportedly aligned to the emotional trajectory. The system is instantiated in an ARPG prototype with generated pixel-art sprites and Unity levels. Evaluation consists of a within-subject user study (n=16) comparing a Baseline episode with an Emotional Arc episode on enjoyment, relevance, and difficulty ratings; a binary arc-identification task; semi-structured interviews; and a post-hoc GoEmotions sentiment analysis of generated arcs. The main reported quantitative result is a significant enjoyability advantage for the Emotional Arc condition (Wilcoxon p=0.016, r=0.789), and 13/16 participants correctly identified the arc condition. Sentiment trajectories show alignment for non-monotonic arcs but not for monotonic Rise/Fall arcs.","tokens_in":10829,"tokens_out":3199,"duration_ms":40687,"significance":"If the central claim is sustained, the framework would be a practical contribution: it gives designers a high-level emotional constraint that can be operationalized by LLMs into playable, branching game content, and it tests the result with both human perception and an independent sentiment classifier. Strengths of the submission include the end-to-end system implementation, the use of an outside emotion classifier rather than the generating LLM itself, and the transparent reporting of qualitative user feedback. However, the significance is currently limited by the small sample, the exploratory design, and an unresolved confound between narrative arc and gameplay difficulty. The paper's value therefore depends on whether the authors can disentangle the emotional arc from difficulty pacing in the only statistically significant result.","major_comments":[{"comment":"The causal attribution of the Enjoyment result (Table 2, p=0.016, r=0.789) to emotional arc integration is undermined by an internal contradiction. The Method states that entity statistics are 'automatically configured based on the emotional arc'; the Case Study reports that enemies in Fall nodes were more likely to have higher attack, health, and ranged capabilities; the Evaluation claims that 'gameplay mechanics, and difficulty parameters were held constant across conditions'; and the Limitations concedes that the Baseline condition 'did not include consistent difficulty modulation' and that ranged-attack enemies may have confounded difficulty comparisons. No per-condition entity statistics are reported. As written, the significant enjoyment difference could reflect difficulty pacing or enemy composition rather than the emotional narrative structure, so the only significant quantitative outcome for RQ1 cannot be causally interpreted. The authors need to report the actual entity-stat distributions per condition and, if those are not available, temper the RQ1 conclusion or provide an additional controlled experiment.","section":"Evaluation (User Study) vs. Method (Entity Generation) and Case Study"},{"comment":"The system-level validation fails for the two primitive labels on which the whole pipeline is built. The paper states that 'monotonic arc types (steady Rise, steady Fall) do not exhibit strong directional trends' and remain flat or moderately positive. Since every canonical arc is composed of Rise and Fall labels and the mind reset prompts condition text on these two labels, the sentiment analysis shows that the basic building blocks are not reliably conveyed to the text. The RQ2 claim that emotional arcs are perceptible to computational emotion models is therefore only supported for non-monotonic composites, not for Rise and Fall themselves. This weakens the central claim that emotional arcs are a practical generative constraint and should be acknowledged explicitly in the abstract and discussion, not only in the Limitations section.","section":"Results (Sentiment Alignment)"},{"comment":"The abstract and conclusion claim that emotional arc integration 'significantly enhances engagement, narrative coherence, and emotional impact.' The only statistically significant quantitative user-rating is Enjoyment; Relevance is not significant (p=0.086), Difficulty is not significant (p=0.573), and no direct Likert or otherwise quantified measure of 'narrative coherence' or 'emotional impact' is reported. The qualitative interview themes are suggestive but cannot carry the weight of a significance claim in the abstract. The claims should be scaled to what the data actually support: a significant enjoyment difference whose attribution is currently confounded, plus qualitative and sentiment-based evidence of perceptibility for non-monotonic arcs.","section":"Abstract, Discussion, and Conclusion"}],"minor_comments":[{"comment":"In the text, bootstrapping is described as using 1,000 samples, but Table 1's caption says 10,000 samples; the discrepancy should be reconciled.","section":"Evaluation (User-Rated Scales)"},{"comment":"Figure 4 would be easier to interpret if the y-axis were explicitly labeled with the valence scale (e.g., the computed V values), since the reader currently has to infer the range and units from the text.","section":"Results (Figure 4)"},{"comment":"Reproducibility would be improved by providing the exact StoryChainRevision prompt templates or an appendix with the full prompts, rather than only the mind reset strings.","section":"Method (Story Generation)"},{"comment":"The Todd et al. reference is incomplete: 'In Proceedings of the 2023 ACM Conference' omits the conference name and page range.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable systems/PCG contribution, but the central empirical claim currently rests on a single small-sample result whose independent variable is internally inconsistent across sections. The authors should be given the chance to resolve the difficulty confound by reporting the actual per-condition entity statistics or by running a follow-up control; if that cannot be done, the abstract and RQ1 conclusion must be substantially weakened. The sentiment result on monotonic arcs is an additional load-bearing weakness that needs either a fix or a frank reframing of RQ2."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a solid systems paper with an honest limitations section, but the central empirical claim is shakier than the abstract admits. The real contribution is the pipeline: Reagan-style Rise/Fall arcs as structural labels for an LLM-generated story DAG, then grounding each node into ARPG entities with difficulty stats. That integrated prototype is new, and the authors cite the component work fairly. They also deserve credit for treating arcs as inputs rather than fitting them to the outcome, and the GoEmotions sentiment check is an independent classifier.\n\nWhat the paper does well: the prototype appears real, the entity schema is clean, the qualitative interview data are useful, and the Limitations section explicitly names several real problems. The 13/16 arc-identification result is a decent perceptibility signal, and the non-monotonic sentiment curves (Rise-Fall-Rise, Fall-Rise-Fall) do show shape alignment. The statistics are simple and appropriate for an exploratory n=16 study.\n\nThe soft spots are real, though. The only statistically significant quantitative result is Enjoyment (p=0.016, n=16), and the evaluation setup makes causal attribution unclear. The Method says entity stats are automatically configured by emotional arc; the Case Study says Fall nodes produced stronger enemies; the Evaluation says difficulty parameters were held constant across conditions; the Limitations concedes the baseline lacked difficulty modulation. Those statements cannot all be true in the way a reader needs them to be. If enemy stats differed between Baseline and Emotional Arc episodes, the enjoyment bump could be difficulty pacing or enemy composition rather than the emotional arc labels. The paper does not report per-condition entity statistics, so the confound cannot be resolved from the text. That is load-bearing for the abstract's claim that arc integration \"significantly enhances engagement, narrative coherence, and emotional impact.\"\n\nThe sentiment analysis is also weaker than the Discussion implies: the monotonic Rise and Fall arcs, which are the primitive building blocks, stay flat and moderately positive. If the simplest arcs do not produce distinguishable trajectories, the emotional-arc operationalization is only partially working. The paper admits this, then later says the trajectories \"closely mirrored\" all six shapes. That overstates it.\n\nBottom line: this deserves peer review. It is a serious, mostly transparent system with a plausible path to a good paper, but not as-is. The authors need to resolve the difficulty-control contradiction, report entity statistics per condition, and recalibrate the claims to what the data actually support. If they do that, it becomes a useful contribution for anyone working on LLM-driven procedural narrative and level generation.","headline":"A genuinely useful LLM+PCG systems paper whose main enjoyment result is undermined by an internal difficulty-control contradiction; worth serious review, but the causal claim needs fixing.","tokens_in":11341,"tokens_out":2683,"would_cite":true,"duration_ms":34134,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Story arcs guide procedural game levels to higher enjoyment by serving as a generative constraint on narrative and difficulty.","keywords":["emotional arcs","procedural content generation","large language models","branching narrative","player experience","sentiment analysis","game level generation","interactive storytelling"],"falsifier":"Generate many stories for each of the six arcs using the exact prompt strings and test whether a strict monotonicity check on per-node valence passes for Rise and Fall; the paper's own data already suggest it would not. A sharper falsifier is to ask blind players to sort monotonic-arc stories into the intended arc order at better-than-chance rates, which would isolate whether the label itself drives the enjoyment effect rather than any generic improvement in prose quality.","tokens_in":10395,"feed_emoji":"🎮","tokens_out":6691,"duration_ms":68790,"temperature":0.7,"pith_summary":"This paper tries to establish that emotional arcs—the rise-and-fall patterns of affective tone that recur in stories—can be used as a working generative constraint for procedural game level generation, not just as a descriptive theory. It builds a pipeline that takes a short user prompt and an arc type (Rise, Fall, or a canonical combination such as Cinderella) and produces a branching story graph whose nodes are labeled Rise or Fall, with each node automatically populated with characters, items, and difficulty-adjusted enemy stats. In a prototype action RPG, the arc-guided episodes scored significantly higher on enjoyment (p=0.016, r=0.789), and 13 of 16 players correctly identified which episode contained an arc. The authors argue this demonstrates that emotional structure can serve as a scaffold for narrative coherence and emotional impact in AI-driven game design.","feed_headline":"Story arcs guide procedural game levels to higher enjoyment","feed_subtitle":"A prototype ARPG shows arc-labeled levels beat flat baselines on enjoyment, and most players can spot the arc.","key_machinery":"The load-bearing object is the pair of primitive emotional labels, Rise and Fall, taken from a six-arc taxonomy of stories. Each node of a directed acyclic story graph receives one of these labels, and the label drives two coupled processes: a 'mind reset' prompt that rewrites the node's prose with an uplifting or somber tone, and automatic configuration of enemy statistics so Fall nodes are harder and more hostile while Rise nodes are friendlier and easier. The graph's edges carry playable triggers ('defeat an enemy', 'talk to an NPC') that map onto dungeon rooms in a Unity prototype, and the whole pipeline is structured as an AI chain that separates story-graph construction, semantic refinement, and entity generation so designers can revise nodes before levels are built.","core_discovery":"The central claim is that embedding a canonical emotional arc into the generation process—by assigning each story node a Rise or Fall label, then using those labels to set the node's textual tone and to scale entity difficulty—produces procedural game narratives that players enjoy more and perceive as more emotionally coherent than affect-neutral baselines. The quantitative anchor is the Enjoyment result: a one-sided Wilcoxon signed-rank test gave p=0.016 with a large effect size (r=0.789), and the bootstrap confidence intervals for Enjoyment did not overlap between conditions. The paper also reports that 13 out of 16 participants (81.25%) correctly identified the arc condition, and that valence trajectories from an independent sentiment classifier matched the intended shapes for non-monotonic arcs such as Rise–Fall–Rise, which the authors take as external validation that the intended narrative structure survives in the generated emotional flow.","pith_inferences":["The framework's reliance on text-heavy narrative likely explains why some players preferred the baseline; a testable extension is to express arc cues through non-textual systems (enemy behavior, lighting, music) and measure whether perceptibility rises while 'too much text' complaints fall.","The monotonic arcs' failure to show directional trends suggests the two prompt strings alone are too weak; an explicit next step is valence supervision per node, such as conditioning generation on target positivity scores.","The positive-valence bias across all arcs hints that both the story generator and the sentiment classifier skew positive; calibrating the classifier on neutral text would clarify whether the flat Rise/Fall trajectories reflect generation or measurement.","Because the evaluation is confined to one ARPG template, the framework's promise of genre transfer remains untested; roguelikes and visual novels would be direct testbeds for the same node-and-edge pipeline."],"forward_implications":["Game developers can use a high-level emotional arc as a generative constraint that improves player enjoyment without hand-authoring every branch of a story.","LLM-based procedural generation can maintain global emotional coherence—not just local textual fluency—when arc labels are injected as explicit prompts.","The same arc label can synchronize narrative tone and gameplay difficulty, since Fall nodes empirically produced stronger enemies and Rise nodes friendlier ones.","Non-monotonic arcs such as Rise–Fall–Rise are reproducible by sentiment analysis, suggesting they are reliable tools for controlled emotional pacing.","The 81.25% identification rate implies arcs are perceptible to players, making them a usable design signal rather than an invisible theoretical property."],"supporting_citations":[{"why":"Supplies the six-basic-shapes taxonomy of emotional arcs and the Rise/Fall primitives that the entire pipeline is built on.","marker":"Reagan et al. 2016"},{"why":"Provides the AI Chaining method that decomposes generation into subtasks with intermediate validation, which the paper credits for controlling narrative structure.","marker":"Wu, Terry, and Cai 2022"},{"why":"Generates branching narrative beats as a directed acyclic graph with a language model; the paper's story-graph generation extends this with emotional arc labels on each node.","marker":"Leandro et al. 2024"},{"why":"Integrates real-time language-model-driven character dialogue with narrative beats, the closest prior model to the paper's entity and dialogue generation.","marker":"Kumaran, Rowe, and Lester 2024"},{"why":"Maps story semantics to spatial layouts and non-player character placements, the basis for translating narrative nodes into dungeon levels.","marker":"Nasir, James, and Togelius 2024"}],"fun_headline_variants":["Emotional arcs guide procedural levels to higher enjoyment","Arc-labeled game levels beat neutral baselines on enjoyment","Procedural narrative arcs raise player engagement in ARPG test","Most players spot the emotional arc in procedural levels","Emotional arc structure improves procedural game narratives"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole pipeline rests on the assumption that the Rise and Fall labels, injected as short tone-setting prompts, reliably produce distinguishable emotional trajectories in the generated text—yet the paper's own sentiment analysis shows that the simplest arcs, steady Rise and steady Fall, come out flat or moderately positive rather than following their intended directions.","fun_headline_variants_meta":{"raw":{"variants":["Emotional arcs guide procedural levels to higher enjoyment","Arc-labeled game levels beat neutral baselines on enjoyment","Procedural narrative arcs raise player engagement in ARPG test","Most players spot the emotional arc in procedural levels","Emotional arc structure improves procedural game narratives"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1562,"prompt_tokens":901,"completion_tokens":661,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":587}},"tokens_in":517,"tokens_out":661,"duration_ms":7604,"temperature":1.0,"reasoning_tokens":587,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T05:08:05.083991+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate many stories for each of the six arcs using the exact prompt strings and test whether a strict monotonicity check on per-node valence passes for Rise and Fall; the paper's own data already suggest it would not. A sharper falsifier is to ask blind players to sort monotonic-arc stories into the intended arc order at better-than-chance rates, which would isolate whether the label itself drives the enjoyment effect rather than any generic improvement in prose quality.","supporting_citations":[{"cited_title":"J.; Mitchell, L.; Kiley, D.; Danforth, C","cited_arxiv_id":null,"evidence_quote":"Supplies the six-basic-shapes taxonomy of emotional arcs and the Rise/Fall primitives that the entire pipeline is built on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Generates branching narrative beats as a directed acyclic graph with a language model; the paper's story-graph generation extends this with emotional arc labels on each node."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Integrates real-time language-model-driven character dialogue with narrative beats, the closest prior model to the paper's entity and dialogue generation."},{"cited_title":"U.; James, S.; and Togelius, J","cited_arxiv_id":null,"evidence_quote":"Maps story semantics to spatial layouts and non-player character placements, the basis for translating narrative nodes into dungeon levels."}],"review_version":1}