{"id":"720881d2-acc3-4e4d-8134-2b2f8b795bf9","arxiv_id":"2505.04890","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A user study of an LLM-powered improv tool for actors finds that unpredictable AI scenarios boost creativity but that overly detailed AI scripts reduce actors' interpretive freedom.","lead":"The paper introduces Theatrical Language Processing, a term for using large language models to support actors' improvisational practice, and evaluates a prototype tool called Scribble.ai with fourteen theater professionals. It reports that actors found AI-generated unpredictable scenarios creatively stimulating but that overly detailed AI scripts limited interpretive freedom.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Causal attribution to AI unpredictability is unsupported because Task 2 lacks a control condition and blinded outcome measures.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the observed effects may be caused by novelty, practice, or demand characteristics rather than by the AI content. My reading of the full text confirms this. The study design section describes Task 2 as a real-time tool usage session followed by interviews and directors' reviews, with no comparison condition. The authors explicitly note that participants were fully briefed on the purpose and procedures, which heightens demand-characteristic risk. Additionally, the 'randomness' input is described in the system but never manipulated in the study, so the specific claim about unpredictability is not tied to a controlled variable. This is a correctness risk rather than an internal inconsistency: the qualitative observations are plausible and consistent with prior creativity-support-tool literature, but the causal language in the abstract and findings goes beyond what the design can support. The recommended check—a three-condition within-subjects experiment with blinded raters—would settle whether the concern lands. Since the reader's verdict is already CONDITIONAL and this concern supports that rather than overturning it, no verdict change is needed.","tokens_in":10074,"tokens_out":3307,"duration_ms":37443,"concrete_test":"Run a preregistered within-subjects study with counterbalanced conditions: (A) improvise with AI-generated irregular scenarios via Scribble.ai, (B) improvise with human-authored irregular scenarios matched for length, novelty, and ambiguity, and (C) free improvisation with no script. Have independent raters blind to condition and hypothesis score video recordings of performances on creativity and problem-solving, and collect participant self-reports after a distractor task that conceals the manipulation. If conditions A and B produce equivalent scores, the causal role of AI unpredictability is not supported; if A exceeds both B and C, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that AI-produced irregular scenarios expanded actors' creativity and heightened their problem-solving skills presupposes a causal effect of the AI content. Task 2 provides no baseline: participants practiced improvisation with Scribble.ai and then reported on their experience, while directors—who knew the tool and the study's purpose—judged actors' adaptability. The protocol explicitly states that participants were 'fully briefed ... on the purpose and procedures' (Study Design), making demand characteristics plausible. There is also no manipulation of the 'randomness' input in the reported study, so the effect of 'unpredictability' is not tied to a measured variable: directors observed actors adapting to 'unconventional topics,' but any novel challenging scenario could produce the same observed adaptation. The negative finding—that overly detailed AI scripts limit interpretive freedom—is better supported by the Task 1 human-vs-AI script comparison, but that comparison uses only three scripts per condition and no matched content analysis. Without a control condition using human-authored irregular scenarios or free improvisation, and without blinded outcome assessment, the study cannot distinguish the tool's effect from practice effects, novelty effects, or social-desirability bias.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a new concept, Theatrical Language Processing (TLP), and an LLM-based creativity support tool, Scribble.ai, which generates improvisational dialogue and monologue scripts from user-provided keywords, genre, and a randomness parameter. The authors report a qualitative user study with 14 participants recruited from a single theater company: 10 actors and 4 directors/acting coaches. In Task 1, participants analyzed three human-authored and three AI-authored scripts to explore perceptions of authorship; in Task 2, participants used Scribble.ai in real-time improvisation and were interviewed afterward, while directors observed their adaptability. The abstract claims that actors expanded their creativity and problem-solving skills when facing AI-produced irregular scenarios, but that AI-generated scripts were often too detailed and limited interpretive freedom and subtext exploration. The paper concludes with proposals for future work on neutral scenes and ambiguity in AI-generated theatrical language.","tokens_in":10362,"tokens_out":4839,"duration_ms":50472,"significance":"If the causal claims were supported, this paper would fill a genuine gap in creativity support tool research, which has largely focused on writers rather than actors, and it would contribute a useful negative result about over-specification in AI-generated scripts. The system description is concrete enough to be reproduced, and the authors are transparent about the exploratory nature of the study and about the participants' own concerns regarding overly directive scripts. The main conceptual contribution, TLP, is essentially a relabeling of prompt-based LLM use for theatrical text generation, so the paper's significance rests on the user study rather than on a new technical method. That study, however, provides only selected quotations and an interpretive summary, with no control condition, no blinded assessment, and no structured qualitative analysis, which severely limits the strength of the causal conclusions drawn in the abstract.","major_comments":[{"comment":"The abstract's causal claim that 'the AI's unpredictability heightened their problem-solving skills' is not supported by the reported design. Task 2 has no control condition, participants were fully briefed on the purpose and procedures, and the directors who evaluated adaptability knew both the tool and the study's hypotheses, so practice effects, novelty effects, and demand characteristics are all plausible alternative explanations. The paper should either reframe the findings as subjective reports of experience or add a comparison condition, such as improvising from human-authored irregular scenarios or no-tool improv, with blinded or at least independent outcome assessment.","section":"Study Design (Task 2) and Findings"},{"comment":"The conclusion that AI-authored scripts are overly detailed and limit interpretive freedom compared with human scripts is based on only three scripts per condition, with no description of how the human scripts were selected, no matching of genre or length, no structured content analysis, and no inter-rater reliability for the reported themes. Please report the full script set, define and quantify categories such as explicit emotion specification and directive stage directions, and demonstrate that the observed difference is systematic rather than idiosyncratic to the six specific scripts used.","section":"Task 1 (Human vs AI-authored scripts)"},{"comment":"The system exposes a 'Creativity Level' or randomness parameter that is central to the claimed effects, but the study never reports what values were used, whether they were varied systematically, or whether participants' behavior differed across levels. Without a manipulation check or logged parametrization, the specific attribution of effects to 'unpredictability' rather than to any novel or challenging scenario content is untestable.","section":"Scribble.ai System Description (Creativity Level)"},{"comment":"All 14 participants came from a single theater company, and the four directors and coaches evaluated their own actors, which introduces dependency between raters and ratees. In addition, the qualitative analysis is not described: there is no interview protocol, no coding scheme, and no procedure for extracting themes, making it difficult to separate the authors' interpretive summary from the participants' actual responses. Please provide the interview protocol, an explicit analysis procedure, and a discussion of how the rater-ratee relationship was handled.","section":"Study Design (Participants) and Findings"}],"minor_comments":[{"comment":"The pseudocode is inconsistent with the prose: the Monologue class references Keyword and Genre rather than OneSentence and Emotion, and the label 'Scriptzing' appears in the interaction flow while the algorithm and body use 'Scriptizing'.","section":"Algorithm 1 and System Description"},{"comment":"References [11], [18], and [33]-[36] are incomplete or non-standard: [11] lacks authors and venue, [18] lists 'Bremen, Germany: ACM Press' as the publisher, and personal communications are presented as references rather than as formative feedback in the text. These should be converted to citable sources or moved to acknowledgements.","section":"References"},{"comment":"The paper uses 'Creativity Level' and 'randomness' interchangeably; please choose one term and define it consistently in the system description, figures, and study narrative.","section":"System Description"},{"comment":"The statement that participants 'erroneously' tied typos to human authorship needs ground-truth clarification, since the paper does not report whether the human-authored scripts actually contained typos.","section":"Findings (Perception of Errors and Vocabulary)"},{"comment":"The captions for Figures 1 and 3 both describe the 'final UI system design'; please verify that each caption matches the displayed component, since one figure is labelled Dialogue and the other Monologue.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an exploratory qualitative HCI study with a small, convenient sample. The central claims in the abstract are stated more strongly than the design can support, and the revision must either substantially temper those claims or add a comparison condition and blinded assessment. The TLP concept is not yet a technical contribution, but the paper could be a useful contribution to the HCI and creativity support tool communities if the evidentiary standards are brought in line with the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: this paper is worth engaging with if you work on creativity support tools, but keep expectations calibrated. It reports a small qualitative study (n=14, drawn from one theater company) of actors improvising with a GPT-3.5 wrapper called Scribble.ai. The genuinely new bit is the user population: prior NLP-for-theater work targets scriptwriters, while this puts improv actors and directors in the foreground and treats interpretive freedom as a design criterion. That is a real gap, and the paper does well by it.\n\nThe strongest result is the negative one: AI scripts that over-specify emotion and action make acting easier but suppress subtext work, and the actors themselves noticed the tradeoff. That finding is convergent across the authorship-judgment task and the live-use task, and it fits prior creativity-support literature. The 'unpredictability helps creativity' finding is plausible but weaker. The stress-test note has it right: Task 2 has no control or baseline, no manipulation of the tool's randomness parameter, participants were fully briefed on the study's purpose, and the directors who judged adaptability knew the tool and the study's aim. So the abstract's causal verbs ('expanded,' 'heightened') outrun what the evidence supports. That said, a control-less qualitative design is common in formative HCI work; the real problem is the framing, and the fix can be either a proper comparison or softer claims.\n\nSoft spots in proportion: the analysis is thin—no coding protocol, mostly selected quotes—which a revision should address. 'Theatrical Language Processing' is a label, not a method. The reference list has incomplete entries and leans on personal communications and demo sessions as design authority; I would ask the authors to fix the citations and rely less on named-expert endorsements. None of this sinks the paper.\n\nAudience: HCI and creativity-support researchers, and people building AI tools for performance domains. It deserves a serious referee rather than a desk reject. I would send it out, with a revision bar that includes reframing the causal claims, reporting a transparent analysis method, and cleaning up the references.","headline":"A modest exploratory HCI study with one honest, useful finding—actors need interpretive room and over-specific AI scripts remove it—wrapped in causal claims the design cannot support.","tokens_in":10762,"tokens_out":5648,"would_cite":true,"duration_ms":54245,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that actors' creativity grows with AI-generated irregular scenarios but falls when AI scripts over-specify emotion.","keywords":["Theatrical Language Processing","Creativity Support Tool","Human-AI Interaction","Large Language Models","Improvisational acting","Scriptwriting","Unpredictability","Subtext"],"falsifier":"A matched experiment could settle the claim: two groups of actors improvise with equally irregular scripts, one set written by the AI tool and one set by humans, with directors rating creativity and adaptability blind. If the human-irregular group improves as much as the AI group, the paper's attribution of the effect to AI unpredictability fails, even though its design recommendation of irregular scenarios would still stand.","tokens_in":9845,"feed_emoji":"🎭","tokens_out":6169,"duration_ms":61691,"temperature":0.7,"pith_summary":"This paper sets out to show that large language models can be shaped into rehearsal partners for actors, and that the value of the AI lies in the irregularity and unpredictability it injects, not in the completeness of the scripts it writes. The authors propose a named area, Theatrical Language Processing (TLP), and build a tool, Scribble.ai, that turns a keyword, a genre, and a randomness level into a dialogue or monologue an actor can improvise with in real time. In a study with ten actors and four directors/acting coaches, they report that unpredictable AI scenarios widened actors' imagination and called on their problem-solving, while AI scripts that explicitly stated emotions or over-specified actions reduced interpretive freedom and blocked subtext exploration. If the finding holds, the design target for AI creativity support in theater is controlled surprise plus deliberate ambiguity.","feed_headline":"AI irregularity stretches actors' improv; AI detail shrinks it","feed_subtitle":"A 14-participant study finds that LLM unpredictability boosts creativity until scripts spell out too much.","key_machinery":"The central mechanism is Scribble.ai's 'Creativity Level,' a user-controlled randomness parameter that determines how much ambiguity and challenge the generated script contains. The generation pipeline first turns the user's keyword, genre, and randomness inputs into a short story abstract, which is then used as the system prompt, allowing the user to add new lines or sudden changes (such as 'introduce a dragon') while the story keeps its central topic. The same pipeline also produces monologues from a single sentence, an emotion, and a randomness level. The paper argues that high randomness is what produces the irregular scenarios that expanded actors' creativity, while the model's tendency to write explicit emotional state labels is what suppresses subtext exploration.","core_discovery":"The paper's central discovery claim is that actors became more creative precisely when the AI-generated scenario was irregular, and that the same AI became a liability when it was too directive. In their own description, participants 'expanded their creativity when faced with AI-produced irregular scenarios,' the AI's 'unpredictability heightened their problem-solving skills' in unfamiliar situations, and scripts that were 'excessively detailed' made performances feel forced and less authentic. The authors call this dual result evidence that AI improvisation support should be engineered for openness—scenarios that create problems for the actor to solve—rather than for ease of execution.","pith_inferences":["A testable extension is to hold the generated scripts fixed and vary only the stated source (AI vs. human); if creativity gains persist in both conditions, the driver is the irregularity of the material rather than the AI's authorship.","The randomness parameter could be treated as an independent variable in a larger study, letting researchers map how subjective creativity ratings and performance quality vary with the degree of script irregularity.","The authors' observation that Scribble.ai writes better stories about objects than about humans suggests a genre of object-centered, non-anthropomorphic monologues as a distinctive niche for AI-assisted theater.","The same over-specification problem likely applies to other AI writing tools: any generator that names emotions instead of implying them may trade clarity for subtext across film, fiction, and game dialogue."],"forward_implications":["Actors and students can use such a tool for solo practice, generating unlimited fresh scenarios without needing a human partner or a prepared prompt bank.","Raising the randomness parameter becomes a deliberate rehearsal technique for training adaptability and interpretation of unfamiliar material.","Improv-oriented AI generators should avoid explicit emotion labels such as 'I'm feeling sad now,' since the study links those labels to reduced interpretive freedom.","Directors and acting teachers can assign AI-generated challenge prompts that fall outside their own habitual writing patterns.","A TLP model trained on neutral or contentless scenes could reintroduce theatrical ambiguity and reclaim the actor's role of building subtext."],"supporting_citations":[{"why":"Defines the creativity support tool concept and its goal of helping more people be more creative, framing the paper's whole purpose.","marker":"[17]"},{"why":"Recommended including directors and acting coaches in the study so that experts familiar with the actors' usual work could judge creativity gains.","marker":"[33]"},{"why":"Provides the foundational improv methods and exercises that the tool's practice scenarios are built on.","marker":"[19]"},{"why":"Establishes improvisation as a core part of actor training, the tradition this work extends.","marker":"[20]"},{"why":"Defines the étude, a short script without an ending that actors must complete, the structural model that AI-generated improv scripts should emulate.","marker":"[25]"},{"why":"Prior work on co-writing screenplays with language models that the paper extends by adding prompt chaining and a user-controlled randomness level.","marker":"[10]"}],"fun_headline_variants":["Actor improv soars with AI chaos, sinks with AI detail","LLM unpredictability sparks acting; over-scripting kills it","Irregular AI boosts improv; detailed AI hinders subtext","For actors, AI's wildness spurs creativity; its specificity stifles"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the measured creativity gains came from the AI-generated content itself, rather than from the novelty of using a new tool or from participants' desire to perform well for the researchers; the study has no control condition without Scribble.ai to rule out those alternatives.","fun_headline_variants_meta":{"raw":{"variants":["Actor improv soars with AI chaos, sinks with AI detail","LLM unpredictability sparks acting; over-scripting kills it","Irregular AI boosts improv; detailed AI hinders subtext","For actors, AI's wildness spurs creativity; its specificity stifles"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000308,"raw_usage":{"total_tokens":1708,"prompt_tokens":840,"completion_tokens":868,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":456,"completion_tokens_details":{"reasoning_tokens":793}},"tokens_in":456,"tokens_out":868,"duration_ms":7592,"temperature":1.0,"reasoning_tokens":793,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:17:53.738444+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A matched experiment could settle the claim: two groups of actors improvise with equally irregular scripts, one set written by the AI tool and one set by humans, with directors rating creativity and adaptability blind. If the human-irregular group improves as much as the AI group, the paper's attribution of the effect to AI unpredictability fails, even though its design recommendation of irregular scenarios would still stand.","supporting_citations":[],"review_version":1}