Pith. sign in

REVIEW 4 major objections 4 minor 31 references

CO-OPERA: A Human-AI Collaborative Playwriting Tool to Support Creative Storytelling for Interdisciplinary Drama Education

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read In CO-OPERA, a conversational tutor expands students' thinking while multi-agent generators turn that thinking into structured script elements; a small usability study indicates the tool keeps students focused on whole logical narrative…

desk verdict A useful design contribution whose central evaluation claim—better narrative focus—is not supported by edit distance and SUS alone; worth reading for the design, worth sending back for a control condition and a real narrative-quality metric. read the letter →

arxiv 2506.00791 v1 pith:QZOE4MHP submitted 2025-06-01 cs.HC

classification cs.HC
keywords human-AIcollaborationplaywritingdramaeducationgenerativeAIstorytellingcreativitysupportsystemusabilitymulti-agent
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that drama-in-education classrooms need a dedicated AI playwriting scaffold because general-purpose chatbots fail at complex educational narratives. It reports need-finding interviews with 13 teachers that surfaced two obstacles: students produce fragmented ideas without structural integration, and generic AI tools give shallow, one-shot outputs. In response, the authors built CO-OPERA, which separates divergent thinking (a conversational tutor that discusses story ideas) from convergent thinking (agents that generate logline, characters, plot, scenes, and dialogue as structured elements). A usability study with 12 middle school students produced a System Usability Scale score of 77.08 and revision data that the authors read as evidence that users engaged critically with AI-generated content and kept whole-narrative logic in view.

What carries the argument

The divergence-convergence interaction mechanism: a conversational tutor engages users in dialogic coaching to expand creative ideas, while functional agents generate standardized playwriting elements (logline, character, plot, scene, dialogue) to converge those ideas. The workflow is staged, so confirmed output at one step becomes structured input to the next agent, scaffolding narrative causality rather than producing a whole script in one pass.

What would settle it

A comparison study in which students write with CO-OPERA, with a generic chatbot, and with no AI tool, with final scripts scored blind by drama teachers on plot coherence, character motivation, and structural completeness, would settle whether the tool actually improves whole-narrative focus; if edit distance is high but blind scores show no difference, the central claim fails.

Watch

Extended reading notes

Core claim

The central claim is that a step-by-step, tutor-plus-agent interaction model can make generative AI usable for educational playwriting, helping students maintain an overall narrative perspective. CO-OPERA operationalizes this as a divergence-convergence loop: at each playwriting stage (logline, characters, plot, scene, dialogue) a conversational tutor broadens the student's thinking through dialogue, and a functional agent then condenses that thinking into structured script elements the user can revise and regenerate. The evaluation evidence is twofold: normalized edit distances between AI-generated elements and students' final versions are taken to show discussion and critical revision, and the SUS score of 77.08 (SD=13.35) is taken as good usability, with the strongest ratings on functional cohesion, which the authors link to enhanced logical structuring for playwriting. The paper concludes that these results confirm CO-OPERA's capability to deliver logic-driven support for playwriting.

Load-bearing premise

The load-bearing premise is that the amount of editing students do to AI-generated text is evidence of critical thinking and whole-narrative development; the paper offers no independent measure of story logic to support that link.

Editorial extensions

If this is right

  • If the claim holds, teachers can deploy CO-OPERA in interdisciplinary drama classes to scaffold playwriting without relying on generic chatbots.
  • Students can iteratively discuss and revise each script element before it is locked into the next stage, giving teachers discussable content at each phase of the creative process.
  • The step-by-step generation makes plot, character, scene, and dialogue separate revisable elements, which supports attention to narrative causality and structural integration.
  • The SUS score of 77.08 suggests the system is usable enough for classroom use, although slower generation speed remains a noted barrier to address.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The edit-distance evidence would be stronger if validated against an independent narrative-quality rubric or a control condition, because larger edits alone do not prove improved story logic.
  • The divergence-convergence design could plausibly transfer to other structured creative writing tasks, such as essays or historical narrative construction, where whole-argument coherence matters.
  • A direct test comparing final scripts written with CO-OPERA, with a generic chatbot, and with no AI tool, scored blind on plot coherence, would settle whether the claimed scaffolding actually changes narrative outcomes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces CO-OPERA, a web-based human-AI collaborative playwriting tool for interdisciplinary drama education. The design is grounded in need-finding interviews with 13 teachers and implements a 'divergence-convergence' interaction mechanism: a conversational tutor supports divergent thinking, while functional agents generate playwriting elements (logline, characters, plots, scenes, dialogues). The evaluation consists of two studies: an edit-distance comparison between AI-generated text and user-modified final versions in a drama class, and a System Usability Scale (SUS) questionnaire with 12 responses from another class. Based on these, the authors claim that CO-OPERA helps users focus on whole logical narrative development during playwriting.

Significance. If the central claim were adequately supported, CO-OPERA would be a valuable contribution to human-AI co-creative writing in educational settings. The paper has concrete strengths: the formative interviews with 13 teachers identify a real instructional need; the system design is clearly described; the authors provide an open repository with playwriting examples and raw data; and the SUS administration follows standard reverse-scoring practice. However, the central claim about improving whole-logical-narrative focus is not directly measured. The evidence consists of edit-distance statistics and self-reported usability, with no independent narrative-quality rubric, no baseline condition, and no control group. The manuscript is therefore best viewed as a system paper with formative and usability findings, and its causal claims need to be substantially tempered or backed by additional analysis.

major comments (4)
  1. [Section 4.1] The central claim that CO-OPERA 'helps users focus on whole logical narrative development' is inferred from normalized edit distance between the AI-generated text and the final user-edited version (Fig. 4(a)-(b)). Edit distance measures textual difference, not narrative logic. A larger edit distance can equally reflect user simplification, topical divergence, or repair of poor AI output; without a rubric for plot causality, character motivation, or structural coherence, and without a baseline condition, this metric cannot support the causal claim. The sentence 'These revisions reflect CO-OPERA's capacity to help users complete discussions and critical thinking throughout playwriting development' is therefore an assertion rather than a demonstrated result.
  2. [Section 4.2] The SUS evaluation provides useful perceived-usability data (77.08, SD=13.35), but SUS is not a measure of narrative coherence or logical structure. The 'Functional Cohesion' subscale (M=3.29) is a self-report item, not an analysis of the resulting scripts, and no control group or alternative tool was used. Consequently, the statement that 'experimental results confirm CO-OPERA's capability to deliver logic-driven support for playwriting' overstates what cross-sectional self-report data can establish.
  3. [Abstract, Section 4, Section 6] The participant population is described inconsistently: the abstract and Section 4.2 describe middle school students, Section 4.1 describes a high school drama class, and the Conclusion states 'we recruited 12 high school students for SUS test.' This inconsistency makes it unclear who participated in each study and how the two evaluations relate to each other, and it prevents the reader from assessing the generalizability of the results. The authors should harmonize the participant descriptions across the abstract, Section 4, and Section 6.
  4. [Section 6 and Abstract] The paper's own Conclusion acknowledges the lack of controlled experiments and the limited sample size. Given that admission, the abstract's causal phrasing ('shows that our CO-OPERA helps users focus on whole logical narrative development') is too strong. Unless a direct narrative-quality measure is added, the wording should be softened to 'suggests' or 'is perceived as supporting' to match the evidence actually reported.
minor comments (4)
  1. [Throughout] There are several typographical and formatting issues: 'Workfloww' in the Section 3.2 heading, 'Futhermore' in the Introduction, 'Visulization' in an author affiliation, 'a online meeting platform' in Section 2.2, and an apparently truncated quote (P5: 'After wasting hours tweaking pro, I could've written this faster myself!') that should be repaired.
  2. [Figure 4] Figure 4 would benefit from a fuller caption and methodology note: the normalization procedure for edit distance, how deletions/insertions are computed, and whether the shown values are per group or per user should be stated. Error bars or per-participant distributions would also help the reader interpret the magnitude of the differences.
  3. [References] Reference [3] (Bangor et al.) is incomplete: the entry lacks full bibliographic details, including year, journal or venue, and volume/page information. Please provide the complete citation.
  4. [Section 4.2] The sentence 'The two or three students in each group operating the system completed SUS questionnaires, resulting in 12 valid responses' is ambiguous about whether the unit of analysis is the student or the group. Please clarify how many groups participated and whether each response corresponds to an individual student or to a group consensus.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper's central claim rests on observed edit-distance and SUS data, not on a parameter fitted from the outcome, so the derivation chain does not reduce to its inputs.

full rationale

The paper contains no formal derivation or predictive chain that could reduce to its own inputs. The central claim that CO-OPERA helps users focus on whole-logical-narrative development is supported by two empirical measurements: normalized edit distance between AI-generated text and the user-edited final version in Section 4.1, and SUS questionnaire responses in Section 4.2. Edit distance is not defined in terms of the target outcome (narrative logic); it is an independent behavioral observation collected from the system logs. The SUS score is compared against an external benchmark (Bangor et al., 2009), so it is not a self-referential validation. No parameter is fitted to a subset of data and then renamed as a prediction, and no uniqueness theorem or self-citation is invoked to forbid alternative explanations. The only self-reference is recruitment from 'our previous study' in Section 2.1, which is not load-bearing for the outcome claim. The paper's evidentiary weakness—asserting that revision distance reflects critical thinking without an independent narrative-quality rubric or a control condition—is a construct-validity limitation, not circularity. The conclusion itself acknowledges the lack of controlled experiments, and that limitation is properly stated. Therefore no circular step is present, and the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No numeric parameters are fitted. The central assumptions are evaluative: the edit-distance proxy, the cognitive-process mapping, the generalizability of the interview sample, and the use of SUS as evidence for narrative outcome. No new physical or theoretical entities are introduced.

assumptions (4)
  • domain assumption Edit distance between AI-generated text and user-edited text is a valid proxy for critical thinking and narrative development.
    Section 4.1 states 'These revisions reflect CO-OPERA's capacity to help users complete discussions and critical thinking'. No validation or independent coding links edit distance to narrative quality.
  • ad hoc to paper The divergence-convergence cognition model maps directly onto the conversational tutor and functional agent components.
    Section 3.2 and Section 5 present this mapping as the theoretical basis; there is no test that the tutor produces divergent thinking or that the agents produce convergence.
  • domain assumption Self-reported SUS scores from 12 students in one class can support the claim that the system improves whole-logical-narrative focus.
    Section 4.2 uses SUS 77.08 and a subscale mean to conclude logic-driven support; SUS measures perceived usability, not learning or narrative quality.
  • domain assumption Interviews with 13 teachers recruited partly through author contacts represent the broader population of drama educators.
    Section 2.1 describes recruitment through a contact list from a previous study and a social media post, which limits representativeness.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CO-OPERA: A Human-AI Collaborative Playwriting Tool to Support Creative Storytelling for Interdisciplinary Drama Education." pith.science (2026). https://pith.science/paper/QZOE4MHP

@misc{pith2026250600791,
  author       = {Pith},
  title        = {Pith review of: CO-OPERA: A Human-AI Collaborative Playwriting Tool to Support Creative Storytelling for Interdisciplinary Drama Education},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QZOE4MHP}},
  note         = {Machine review of arXiv:2506.00791}
}
read the original abstract

Drama-in-education is an interdisciplinary instructional approach that integrates subjects such as language, history, and psychology. Its core component is playwriting. Based on need-finding interviews of 13 teachers, we found that current general-purpose AI tools cannot effectively assist teachers and students during playwriting. Therefore, we propose CO-OPERA - a collaborative playwriting tool integrating generative artificial intelligence capabilities. In CO-OPERA, users can both expand their thinking through discussions with a tutor and converge their thinking by operating agents to generate script elements. Additionally, the system allows for iterative modifications and regenerations based on user requirements. A system usability test conducted with middle school students shows that our CO-OPERA helps users focus on whole logical narrative development during playwriting. Our playwriting examples and raw data for qualitative and quantitative analysis are available at https://github.com/daisyinb612/CO-OPERA.

Figures

Figures reproduced from arXiv: 2506.00791 by the authors.

Figure 1
Figure 1. Overview: (a) Human-AI co-creation framework; (b) Playwriting examples with dramatic elements [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. (a)Human-AI interaction framework of playwriting; (b) Interdisciplinary drama class through teamwork [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Multi-agent workflow: dataflow among functional agents, assistant tutors, and users. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: (a) Absolute edit distance; (b) Relative edit distance; (c) SUS questionnaire results:Willingness (Mean=3.04; Q1=3.58, Q2=2.50), [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

31 extracted references · 8 canonical work pages

  1. [1]

    Isabel Machado Alexandre, David Jardim, and Pedro Faria Lopes. [n. d.]. Maths4Kids: Telling Stories with Maths. InProceedings of the Intelligent Narrative Technologies III Workshop(New York, NY, USA, 2010-06-18)(INT3 ’10). Association for Computing Machinery, 1–6. https://doi.org/10. 1145/1822309.1822313

  2. [2]

    Yuichiro Anzai and Herbert A Simon. 1979. The theory of learning by doing.Psychological review86, 2 (1979), 124

  3. [3]

    Aaron Bangor. [n. d.]. Determining What Individual SUS Scores Mean: Adding an Adjective Rating Scale. 4, 3 ([n. d.])

  4. [4]

    Guillermo Bernal. [n. d.].Paper Dreams: Real-Time Human and Machine Collaboration for Visual Story Development. MIT Media Lab. https: //www.media.mit.edu/publications/real-time-human-and-machine-collaboration-for-visual-story-development/

  5. [5]

    David Wallace Booth and Kathleen Gallagher. [n. d.].How Theatre Educates: Convergences and Counterpoints with Artists, Scholars and Advocates. University of Toronto Press. googlebooks:rRhk4YA8Rz4C

  6. [6]

    Jiaju Chen, Minglong Tang, Yuxuan Lu, Bingsheng Yao, Elissa Fan, Xiaojuan Ma, Ying Xu, Dakuo Wang, Yuling Sun, and Liang He. 2025. Characterizing LLM-Empowered Personalized Story-Reading and Interaction for Children: Insights from Multi-Stakeholder Perspectives. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems(Yokohama, Japan...

  7. [7]

    Liuqing Chen, Shuhong Xiao, Yunnong Chen, Ruoyu Wu, Yaxuan Song, and Lingyun Sun. 2024. ChatScratch: An AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12. https://doi.org/10.1145/3613904.3642229 arXiv:2402.04975 [cs]

  8. [8]

    Nicholas Davis, Holger Winnemöller, Mira Dontcheva, and Ellen Yi-Luen Do. 2013. Toward a cognitive theory of creativity support. InProceedings of the 9th ACM Conference on Creativity & Cognition. 13–22

Show all 31 references
  1. [10]

    Bederson, and Alex Quinn

    Allison Druin, Benjamin B. Bederson, and Alex Quinn. [n. d.]. Designing Intergenerational Mobile Storytelling. InProceedings of the 8th International Conference on Interaction Design and Children(New York, NY, USA, 2009-06-03)(IDC ’09). Association for Computing Machinery, 325...

  2. [11]

    Min Fan, Xinyue Cui, Jing Hao, Renxuan Ye, Wanqing Ma, Xin Tong, and Meng Li. 2024. StoryPrompt: Exploring the Design Space of an AI- Empowered Creative Storytelling System for Elementary Children. InExtended Abstracts of the CHI Conference on Human Factors in Computing System...

  3. [12]

    Min Fan, Xinyue Cui, Wanqing Ma, Haiyan Li, Xin Tong, Lin Yang, and Yonghui Wang. 2025. From Words to Wonder: Designing and Evaluating an AI-Empowered Creative Storytelling System for Elementary Children. InProceedings of the 2025 CHI Conference on Human Factors in Computing S...

  4. [13]

    Arzu Guneysu, Helena Isabel Silva Reis, Sanna Kuoppamäki, and Cristina Maria Sylla. 2024. Enhancing Autism Therapy through Smart Tangible- Based Digital Storytelling: Co-Design of Activities and Feasibility Study. InProceedings of the 23rd Annual ACM Interaction Design and Chi...

  5. [14]

    Jiahao Guo, Yuyu Lin, Hongyu Yang, Junwu Wang, Shuo Li, Enmao Liu, Cheng Yao, and Fangtian Ying. 2020. Comparing the Tangible Tutorial System and the Human Teacher in Intangible Cultural Heritage Education. InProceedings of the 2020 ACM Designing Interactive Systems Conference...

  6. [15]

    Ariel Han and Zhenyao Cai. [n. d.]. Design Implications of Generative AI Systems for Visual Storytelling for Young Learners. InProceedings of the 22nd Annual ACM Interaction Design and Children Conference(New York, NY, USA, 2023-06-19)(IDC ’23). Association for Computing Machi...

  7. [16]

    Senyu Han, Lu Chen, Li-Min Lin, Zhengshan Xu, and Kai Yu. 2024. IBSEN: Director-Actor Agent Collaboration for Controllable and Interactive Drama Script Generation. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)...

  8. [17]

    Shih, and Kyungsik Han

    Youngseung Jeon, Seungwan Jin, Patrick C. Shih, and Kyungsik Han. 2021. FashionQ: An AI-Driven Creativity Support Tool for Facilitating Ideation in Fashion Design. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Computi...

  9. [18]

    Jacqueline Kory and Cynthia Breazeal. [n. d.]. Storytelling with Robots: Learning Companions for Preschool Children’s Language Development. InThe 23rd IEEE International Symposium on Robot and Human Interactive Communication(2014-08). 643–648. https://doi.org/10.1109/ROMAN.201...

  10. [19]

    1991.Situated learning: Legitimate peripheral participation

    Jean Lave and Etienne Wenger. 1991.Situated learning: Legitimate peripheral participation. Cambridge university press

  11. [20]

    Yuyu Lin, Jiahao Guo, Yang Chen, Cheng Yao, and Fangtian Ying. 2020. It Is Your Turn: Collaborative Ideation with a Co-Creative Robot through Sketch. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems (CHI ’20). Association for Computing Machinery, ...

  12. [21]

    When He Feels Cold, He Goes to the Seahorse

    Di Liu, Hanqing Zhou, and Pengcheng An. 2024. "When He Feels Cold, He Goes to the Seahorse"-Blending Generative AI into Multimaterial Storymaking for Family Expressive Arts Therapy. InProceedings of the CHI Conference on Human Factors in Computing Systems. 1–21. https: //doi.o...

  13. [22]

    Mathewson, Jaylen Pittman, and Richard Evans

    Piotr Mirowski, Kory W. Mathewson, Jaylen Pittman, and Richard Evans. 2023. Co-Writing Screenplays and Theatre Scripts with Language Models: Evaluation by Industry Professionals. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. ACM, Hamburg Germa...

  14. [23]

    Alisha Panjwani. [n. d.]. Constructing Meaning: Designing Powerful Story-Making Explorations for Children to Express with Tangible Computational Media. InProceedings of the 2017 Conference on Interaction Design and Children(New York, NY, USA, 2017-06-27)(IDC ’17). Association ...

  15. [24]

    Kimiko Ryokai, Hayes Raffle, and Robert Kowalski. [n. d.]. StoryFaces: Pretend-Play with Ebooks to Support Social-Emotional Storytelling. In Proceedings of the 11th International Conference on Interaction Design and Children(New York, NY, USA, 2012-06-12)(IDC ’12). Association...

  16. [25]

    Michael Schlauch, Cristina Sylla, and Maitê Gil. [n. d.]. Investigating Social Emotional Learning at Primary School through Guided Interactive Storytelling. InExtended Abstracts of the 2022 Annual Symposium on Computer-Human Interaction in Play(New York, NY, USA, 2022-11-07)(C...

  17. [26]

    Zizhen Wang, Jiangyu Pan, Duola Jin, Jingao Zhang, Jiacheng Cao, Chao Zhang, Zejian Li, Preben Hansen, Yijun Zhao, Shouqian Sun, and Xianyue Qiao. 2025. CharacterCritique: Supporting Children’s Development of Critical Thinking through Multi-Agent Interaction in Story Reading. ...

  18. [27]

    Weiqi Wu, Hongqiu Wu, Lai Jiang, Xingyuan Liu, Hai Zhao, and Min Zhang. [n. d.]. From Role-Play to Drama-Interaction: An LLM Solution. In Findings of the Association for Computational Linguistics: ACL 2024(Bangkok, Thailand, 2024-08), Lun-Wei Ku, Andre Martins, and Vivek Sriku...

  19. [28]

    Wenjie Xu, Jiayi Ma, Jiayu Yao, Weijia Lin, Chao Zhang, Xuanhe Xia, Nan Zhuang, Shitong Weng, Xiaoqian Xie, Shuyue Feng, Fangtian Ying, Preben Hansen, and Cheng Yao. 2023. MathKingdom: Teaching Children Mathematical Language Through Speaking at Home via a Voice- Guided Game. I...

  20. [29]

    Lili Yao, Nanyun Peng, Ralph Weischedel, Kevin Knight, Dongyan Zhao, and Rui Yan. 2019. Plan-and-Write: Towards Better Automatic Storytelling. Proceedings of the AAAI Conference on Artificial Intelligence33, 01 (July 2019), 7378–7385. https://doi.org/10.1609/aaai.v33i01.33017378

  21. [30]

    Chao Zhang, Xuechen Liu, Katherine Ziska, Soobin Jeon, Chi-Lin Yu, and Ying Xu. 2024. Mathemyths: Leveraging Large Language Models to Teach Mathematical Language through Child-AI Co-Creative Storytelling. https://doi.org/10.1145/3613904.3642647 arXiv:2402.01927 [cs]

  22. [31]

    Chao Zhang, Cheng Yao, Jiayi Wu, Weijia Lin, Lijuan Liu, Ge Yan, and Fangtian Ying. [n. d.]. StoryDrawer: A Child–AI Collaborative Drawing System to Support Children’s Creative Visual Storytelling. InCHI Conference on Human Factors in Computing Systems(New Orleans LA USA, 2022...

  23. [32]

    Zheng Zhang, Ying Xu, Yanhao Wang, Bingsheng Yao, Daniel Ritchie, Tongshuang Wu, Mo Yu, Dakuo Wang, and Toby Jia-Jun Li. 2022. StoryBuddy: A Human-AI Collaborative Chatbot for Parent-Child Interactive Storytelling with Flexible Parental Involvement. InProceedings of the 2022 C...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.