{"id":"8322b054-30f8-4d6d-a1fc-d493c13a4fa4","arxiv_id":"2501.08474","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review arguing that live improvised comedy should be the central testbed for evaluating computational humor systems.","lead":"This position paper reviews AI comedy systems used in live theater and argues that performing in front of real audiences is the best way to test computational humor. It groups the challenges into embodiment, timing and interaction, and human interpretation of AI output.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Live-performance evaluation is confounded by human curation and performance, so the paper supports testing human-AI co-creative systems rather than AI humor alone.","rationale":"The reader's weakest assumption correctly identifies audience response as confounded by novelty and human performer skill. My concern is closely related but more internal to the paper's own evidence: the flagship Improbotics pipeline includes explicit human-in-the-loop curation, and the paper concedes that audience evaluation is often focused on the human performers. This means that even if audience response is valid and reliable, it measures a human-AI co-creative system rather than the computational humor system alone. The reader's version emphasizes external confounds; mine emphasizes that the paper's own system descriptions prevent attribution to the AI. This strengthens the case for a conditional acceptance: the review is useful and the examples are real, but the abstract and Section 5 overclaim by asserting live performance is the ideal substrate for evaluating computational humor systems. The recommendation could be addressed by reframing the claim to 'live performance is a valuable testbed for human-AI co-creative comedy systems' and by adding either an attribution analysis of existing data or a controlled comparison separating curation from generation. The paper does include honest caveats in Sections 2.1, 3.1, and 3.2, and it cites independent work beyond the authors' own, which supports its value as a position statement even though the central claim needs qualification. Therefore the existing CONDITIONAL verdict remains appropriate.","tokens_in":9180,"tokens_out":3368,"duration_ms":38529,"concrete_test":"Conduct a controlled live experiment, or re-analyze the survey data from Branch et al. (2024), comparing three conditions with the same AI-generated lines and the same human performer: (a) human-curated AI lines as in Improbotics, (b) randomly selected AI lines with no human curation, and (c) human-written lines matched for length and topic, delivered in counterbalanced order to comparable audiences. Measure laughter episodes, audience ratings of funniness, and post-show surveys. If condition (a) is indistinguishable from (c) while condition (b) is significantly worse, the live setting's success is driven by human curation and performance rather than by the LLM, and the paper's central claim must be narrowed to co-creative systems. If condition (b) matches (a) and (c), the attribution concern would not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Abstract, Section 5) is that live audiences provide a uniquely realistic testbed for computational humor systems. For that claim to hold, audience laughter or engagement must be attributable to the AI's comedic quality. The paper's own flagship systems violate this attribution requirement. Section 3.1 describes Improbotics relying on human-in-the-loop curation, a 'writer's room' that selects AI-generated lines, and in 2024 the Cyborg performer simultaneously curates lines while delivering them. Section 4 states that human actors attempt to justify 'seemingly absurd' AI-generated text, and Section 3.2 concedes that when AI is used as a writing tool, 'audience evaluation is focused primarily on the human performers.' Section 2.1 explicitly hypothesizes that audiences may have evaluated the novelty of show premises in addition to comedic quality. Thus none of the cited live successes isolates the AI system's contribution to humor. The evidence supports live theater as a testbed for human-AI co-creative systems, not for computational humor systems in isolation. This is an internal tension rather than a disagreement with external consensus: the paper's own described pipelines and caveats undermine the abstract's stronger claim. The manuscript also reports no failed shows or systematic negative results, further weakening the inference from selected successes to a general methodological prescription.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper reviews recent work at the intersection of computational humor generation and live performance, drawing mostly on the authors' own productions (Improbotics, HumanMachine, Dramatron-related shows) and a handful of external projects. It argues that AI comedy systems should be evaluated live, in front of audiences, under real-time constraints, and that improvised comedy is the ideal substrate for such evaluation. The paper organizes the field into three challenge areas: embodiment and anthropomorphism, timing and audience interaction, and human interpretation of AI-generated absurdity. It then discusses how live-audience constraints should shape evaluation methodology.","tokens_in":9550,"tokens_out":3000,"duration_ms":33145,"significance":"If the central claim were established, the paper would shift the evaluation of computational humor away from crowdsourced rating tasks toward ecologically valid live settings, which is a genuinely useful corrective given known problems with decontextualized text evaluation. The paper also usefully compiles a scattered set of artistic and technical works and articulates three distinct challenge areas that are likely to be relevant for any future deployable system. The main contribution is conceptual framing rather than empirical proof: the manuscript offers no systematic comparison, no failed cases, and no protocol for attributing audience response to the AI system rather than to human performers or the novelty of the premise.","major_comments":[{"comment":"The paper's central claim that live audiences provide a testbed for \"computational humor systems\" is contradicted by its own descriptions of the flagship systems. Section 3.1 states that Improbotics relies on human-in-the-loop curation via a \"writer's room\" that selects AI-generated lines, and that in 2024 the Cyborg performer curates lines while delivering them; Section 3.2 concedes that when AI is used as a writing tool, \"audience evaluation is focused primarily on the human performers\"; and Section 2.1 hypothesizes that audiences may have evaluated the novelty of show premises rather than comedic quality. Under these conditions, audience laughter or engagement cannot be attributed specifically to the AI's comedic output. The claim should be narrowed to human-AI co-creative systems, or the paper should specify an attribution methodology capable of isolating the AI's contribution.","section":"Abstract, §2.1, §3.1, §3.2"},{"comment":"The paper explicitly selects \"examples of successful AI-infused shows\" (Abstract) and does not report any failed shows, negative results, or systematic sampling criteria. This makes it impossible to infer the prescriptive claim that live performance is the \"perfect substrate\" for evaluating computational humor. The evidence presented supports only the weaker statement that live performance is one possible testbed among several. A position paper can reasonably advance a methodological thesis, but it should at least address selection effects and discuss what a disconfirming observation would look like.","section":"Abstract and §5"},{"comment":"The evaluation toolbox listed in Section 5—audience surveys, laughter microphones, focus groups, and Creativity Support Index metrics—is presented as applicable to live settings, but the paper never explains how any of these tools would separate the AI system's contribution from the human performer's delivery, the human curator's line selection, or the audience's reaction to the premise. This missing protocol is load-bearing because it is the only route from anecdotal reports to the paper's methodological thesis. Adding a concrete evaluation design with confound controls would considerably strengthen the argument.","section":"§5"}],"minor_comments":[{"comment":"The sentence \"we hypothesize that audiences may have evaluated the novelty of the premises of those shows in addition to their comedic quality\" is a significant caveat that is never revisited in the discussion; the paper would benefit from proposing a way to test this hypothesis empirically.","section":"§2.1"},{"comment":"There are several typographical and phrasing issues: \"Incidently\" should be \"Incidentally\" (Section 4), and \"it thus provides with a realistic\" (Section 1) should be \"it thus provides a realistic.\" These should be corrected.","section":"§1, §4"},{"comment":"Many footnotes (e.g., Footnotes 2–16) are bare URLs without access dates or archival identifiers. For a review paper, accessible and citable sources would improve reproducibility.","section":"References/Footnotes"},{"comment":"The reference to Sutskever et al. (2014) for \"large language models\" is inaccurate: that work introduced sequence-to-sequence learning, not what is now called an LLM. Consider citing a contemporary neural conversational model (e.g., Vinyals and Le, 2015) more precisely.","section":"§2.1 and References"},{"comment":"The paper is titled a \"Review\" but provides no explicit inclusion criteria or scope statement. A brief paragraph describing how the corpus of shows was assembled would help readers understand the intended coverage.","section":"Title and §1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is essentially a position/review paper grounded in the authors' own artistic practice. I do not see a fatal internal inconsistency, but the load-bearing inference from selected successes to a general methodological prescription needs careful reframing. The authors have an opportunity to make the paper stronger by narrowing the claim to co-creative systems and adding a concrete evaluation protocol with explicit confound controls. I would also flag that the heavy reliance on the authors' own productions might be perceived as a conflict-of-interest disclosure concern, though the paper does identify the affiliation openly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a review and position statement, not a new result. What it does well: it is the most complete catalog I have seen of live comedy LLM systems, from Knight's Nao robot to Improbotics, Comedians vs. AI, and THEaiTRE. The three-challenge taxonomy (embodiment, timing, and human interpretation of absurd AI output) is a genuinely useful organizing frame, and the paper gives credit to external work by Winters, Fitter, and others. It also openly concedes several confounds: Section 2.1 hypothesizes that audiences may be reacting to novelty, Section 3.2 admits that when AI is a writing tool, evaluation focuses on the human performers, and Section 3.1 describes human-in-the-loop line curation. That self-awareness is real, not performative.\n\nThe soft spots are the ones flagged in the stress test. The abstract claims live performance is the 'perfect substrate' for assessing computational humor systems, but none of the cited successes isolates the AI system's contribution. Improbotics has a human curator selecting lines; actors are instructed to justify absurd text; the Cyborg performer reads from AR glasses while maintaining eye contact. So audience laughter is attributable to the whole human-AI co-creative system, not to the LLM alone. That is an internal tension with the paper's own descriptions, and it is the load-bearing flaw in the strongest version of the argument. The absence of any failed shows or negative results makes the inference from selected successes to a general methodological prescription even shakier.\n\nThat said, the paper does not actually need the strongest version. Its Section 5 pivots to framing AI as a creativity support tool and evaluation as feedback on the human-AI collaboration. That framing is supported by the evidence they present and by the external literature they cite. If the authors toned the abstract and discussion to match that framing, the paper would be a solid contribution rather than an overclaim.\n\nWho is this for? Researchers in computational humor evaluation, HCI people studying co-creative systems, and digital performance scholars. It is worth a serious referee: the survey alone justifies referee time, and the position is debatable in a productive way. I would send it to review with a request to revise the central claim and add a section on failed or mixed shows. It deserves a conditional accept, not a desk reject.","headline":"A useful survey and a defensible position that overreaches in its abstract: the live-performance examples do not isolate the AI's comedic contribution, but the paper's own creativity-support framing is more honest and worth engaging.","tokens_in":9903,"tokens_out":1280,"would_cite":true,"duration_ms":16228,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that computational humor should be judged in live performance, in front of real audiences, under real-time constraints, with improvised comedy as the ideal test substrate.","keywords":["computational humor","live performance","improvisational theater","large language models","human-AI collaboration","comedy evaluation","audience feedback","comedy Turing test"],"falsifier":"Take a fixed set of AI-written jokes and have the same comedian deliver them to a live audience while a separate group rates the same jokes as plain text, with independent judges scoring both. If live laughter and ratings track the performer's timing and delivery rather than the joke text, the claim that live performance cleanly measures the AI's comedic quality loses support.","tokens_in":9008,"feed_emoji":"🎭","tokens_out":6918,"duration_ms":60307,"temperature":0.7,"pith_summary":"This position paper argues that the best way to know whether an AI system is funny is to put it on stage in front of a live audience and let it work under real-time pressure. It reviews recent AI-comedy performances—robots delivering jokes, human improvisers fed lines by language models, adversarial human-versus-AI shows, and virtual-reality performances—and identifies three recurring challenges: machine embodiment and anthropomorphism, comedic timing and audience interaction, and making sense of absurd AI-generated text. The authors claim that these challenges make live improvised comedy an ideal testbed, because paying audiences bring realistic expectations, performers can fold real-time feedback into the show, and the improv tradition of 'celebrating failure' creates a safe environment for testing AI. If the paper is right, computational humor research should shift from static text and crowdsourced ratings toward live performance, and AI comedy tools should be designed as creativity-support collaborators for human comedians rather than autonomous joke machines.","feed_headline":"Live theater is the right testbed for AI comedy","feed_subtitle":"A position paper says real audiences and real-time pressure reveal what joke-generating systems can actually do","key_machinery":"The paper's central object is the live improvisational comedy show staged as an experimental setting, and its concrete working unit is the 'Cyborg' performer: a human actor who receives real-time AI-generated lines through an earpiece or augmented-reality glasses and must deliver them with comedic timing to a live audience. This setup couples a language model to a human body and to an audience, which forces the three challenges the paper discusses—embodiment, timing, and interpretation—into view at once. In many of the shows reviewed, a second human 'writer's room' curator selects which AI lines reach the performer, adding another layer of human interpretation on top of the model output.","core_discovery":"The central claim is that AI comedy should be evaluated live, in front of audiences sharing physical or online spaces, under real-time constraints, and that improvised comedy is the perfect substrate for deploying and assessing computational humor systems. The paper grounds this claim in a survey of AI-infused shows, from robot standup and human-machine improv troupes to adversarial comedy battles and deep-fake games, and organizes the open questions into three groups: embodiment and anthropomorphism, comedic timing and audience interaction, and human interpretation of seemingly absurd AI output. It then argues that these live constraints reshape methodology: any evaluation method must work around real audiences and performance spaces, and it concludes that the right relationship between comedians and AI tools is collaborative, positioning the AI as a creativity support tool rather than a replacement performer.","pith_inferences":["If live performance becomes the standard testbed, the field would benefit from a portable measurement kit—laughter sensing, performer surveys, and audience polls—so different shows can be compared across venues.","The same 'safe failure' culture of improv could make live co-creation a general test for other real-time generative AI, such as music improvisation or interactive game dialogue, not just humor.","Adopting the paper's framing would imply that decontextualized joke datasets become less decisive for progress, because the timing and social context that matter most are absent from them."],"forward_implications":["Evaluation of AI humor should move from static text ratings and crowdsourced surveys toward live performances with real-time audience feedback.","Metrics for AI comedy must include timing, audience engagement, and performer experience, not just whether the words are funny.","Computational humor systems are best treated as creativity support tools for human comedians, not as autonomous comedians.","The comedy Turing test will not cleanly separate human from machine as language models improve, because human performers can imitate AI and fool audiences."],"supporting_citations":[{"why":"Supplies the founding premise that humor is a fundamental challenge for computational systems.","marker":"Raskin, 1979"},{"why":"Defines improvisational theater's principles and the 'safe failure' culture that makes the stage a testbed.","marker":"Johnstone, 1979"},{"why":"Characterizes theatrical improvisation as real-time dynamic problem solving, supporting the substrate claim.","marker":"Magerko et al., 2009"},{"why":"Documents the Improbotics Cyborg setup and the timing constraints live performance places on AI systems.","marker":"Mathewson and Mirowski, 2018"},{"why":"Early robot standup that uses audience laughter to adjust joke selection, the prototype for live audience measurement.","marker":"Knight et al., 2011"},{"why":"Introduces the comedy Turing test and the concealed-cyborg performances used to evaluate audience perception of AI.","marker":"Mathewson and Mirowski, 2017b"},{"why":"Reports design and evaluation of dialogue LLMs for co-creative improvised theater, including surveys and curation methods.","marker":"Branch et al., 2024"},{"why":"Proposes focus groups and Creativity Support Tool metrics with comedians as a live-compatible way to assess AI humor.","marker":"Mirowski et al., 2024"},{"why":"Represents crowdsourced evaluation of AI-generated jokes that the paper contrasts with live-performance evaluation.","marker":"Gorenz and Schwarz, 2024"},{"why":"Supplies the critique of crowdsourced open-ended text evaluation that motivates preferring live audiences.","marker":"Karpinska et al., 2021"}],"fun_headline_variants":["AI comedy needs live audiences to be truly tested","Improvised comedy is the proving ground for AI humor","Live stage reveals what AI jokes can really do","Real-time constraints make live comedy the AI test","Theater is the laboratory for AI comedy evaluation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that a live audience's response measures the AI system's comedic quality, even though the paper itself concedes that audiences may be reacting to the novelty of the show's premise or to the human performer's delivery.","fun_headline_variants_meta":{"raw":{"variants":["AI comedy needs live audiences to be truly tested","Improvised comedy is the proving ground for AI humor","Live stage reveals what AI jokes can really do","Real-time constraints make live comedy the AI test","Theater is the laboratory for AI comedy evaluation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000597,"raw_usage":{"total_tokens":2758,"prompt_tokens":875,"completion_tokens":1883,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":1811}},"tokens_in":491,"tokens_out":1883,"duration_ms":12866,"temperature":1.0,"reasoning_tokens":1811,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:24:32.834056+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fixed set of AI-written jokes and have the same comedian deliver them to a live audience while a separate group rates the same jokes as plain text, with independent judges scoring both. If live laughter and ratings track the performer's timing and delivery rather than the joke text, the claim that live performance cleanly measures the AI's comedic quality loses support.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the founding premise that humor is a fundamental challenge for computational systems."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines improvisational theater's principles and the 'safe failure' culture that makes the stage a testbed."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Characterizes theatrical improvisation as real-time dynamic problem solving, supporting the substrate claim."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the Improbotics Cyborg setup and the timing constraints live performance places on AI systems."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Early robot standup that uses audience laughter to adjust joke selection, the prototype for live audience measurement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports design and evaluation of dialogue LLMs for co-creative improvised theater, including surveys and curation methods."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the critique of crowdsourced open-ended text evaluation that motivates preferring live audiences."}],"review_version":1}