{"id":"35b48420-6d63-4e28-acf2-b8c8b94e46d7","arxiv_id":"2503.15498","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A performance paper reports on Revival, an improvised audiovisual concert where human musicians collaborate with AI agents trained on small curated datasets, but provides no evaluation of the claimed creative outcomes.","lead":"This paper describes Revival, a live audiovisual performance by the K-Phi-A collective that pairs human percussion and electronics with AI musical agents and an AI-driven visual synthesizer. It argues that such human-AI co-creation can work in improvisational music and visual art.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No recording or documentation verifies that Revival was actually performed; the central 'showcases real-time human-AI co-creation' claim is unsupported, and Appendix A.1 says 'we are prepared to perform if needed.'","rationale":"The reader's formal weakest_assumption focuses on the 55-dimensional feature vectors and FFT parameters. That is a real but secondary technical risk: even if these features are imperfect, the artwork could still demonstrate the potential of human-AI co-creation if a compelling performance is documented. The truly load-bearing premise is more basic: Revival must have been performed, and the agents' behavior in that performance must have been causally tied to the human input. The manuscript does not establish this. Its own appendix says 'we are prepared to perform if needed,' and no video or audio artifact is included or concretely verified beyond a URL. Without such an artifact, claims of 'dynamic response' and 'emulation of complex musical styles' cannot be adjudicated. This is not an accusation of misconduct; it is an assessment of what evidence the text provides. The reader's rationale already notes the absence of recordings, code, data, and user studies, but the formal weakest_assumption is narrower. My concern thus partially agrees with the reader while relocating the load-bearing issue from feature-extraction quality to the unverified existence and documented behavior of the performance itself. The proposed test—checking for a full recording and, if present, measuring causal responsiveness between human input and agent output—would settle this concern concretely. Until that evidence is available, the central claim should be treated as unverified rather than conditionally accepted on the basis of the system description alone.","tokens_in":6280,"tokens_out":4264,"duration_ms":50794,"concrete_test":"Open the linked project page (https://www.metacreation.net/projects/revival-art) and confirm whether a complete, unedited recording of Revival exists. If no such recording exists, the paper should be reclassified as a proposal or practice note, not a documented demonstration of real-time co-creation. If a recording exists, extract the separate output stems for MASOM and SpireMuse and cross-correlate their onsets with the human percussion and electronic input events; a causal, musically plausible response (e.g., median latency under roughly one second and significantly above chance in a permutation test) would support the 'dynamically respond' claim.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central claim—that Revival 'showcases the potential of AI and human collaboration in improvisational artistic creation'—requires that a real performance actually occurred and that the AI agents responded causally to the human performers. The paper provides no recording, no time-aligned audio stems, no code, no performance log, and no audience study. The weakest point is not the 55-dimensional feature vector or the FFT window; it is the unverified existence of the performance itself. Appendix A.1 states 'we are prepared to perform if needed,' which reads as a proposal rather than documentation of a realized artwork. Section 3 is a general literature review and contributes no direct evidence. The abstract's 'features real-time co-creative improvisation' and the conclusion's 'performance demonstrates' therefore rest on an assumption that is neither internally established nor externally checkable from the manuscript. This is a missing-evidence concern, not a claim that the work is fraudulent; a published recording or an independent technical rider would resolve it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes 'Revival,' a live audiovisual performance by the artist collective K-Phi-A, in which a human percussionist and an electronic musician improvise alongside AI musical agents (MASOM and SpireMuse) and an AI-driven visual synthesizer (Autolume). The authors claim that the agents, trained on a small curated corpus of works by deceased composers and the collective's own compositions, dynamically respond to human input in real time and emulate complex musical styles, thereby showcasing human-AI co-creation. The manuscript provides a system description, a review of research-creation methodology, a discussion of challenges, and an appendix with a technical rider and stage plot. No recording, user study, performance log, code, or other direct documentation of the claimed performance is included.","tokens_in":6601,"tokens_out":3381,"duration_ms":37734,"significance":"If the central claim were substantiated, the work would offer a valuable example of real-time human-AI co-creation in musical performance using ethically motivated small-data training. The paper's strengths include the integration of previously developed systems (MASOM, SpireMuse, Autolume), a clear small-data rationale, and a detailed technical setup. However, as submitted, the paper functions more as a performance proposal or system description than as a verified research-creation study: the load-bearing claim that 'Revival showcases the potential of AI and human collaboration' is asserted rather than demonstrated, and no evidence of an actual realized performance is provided.","major_comments":[{"comment":"The abstract and conclusion state that the 'performance demonstrates' and 'features real-time co-creative improvisation,' but Appendix A.1 says 'we are prepared to perform if needed' and describes the piece as 'designed to last approximately 30 minutes,' which reads as a proposal rather than documentation of a realized work. The manuscript provides no recording, time-aligned audio stems, performance log, or audience study. This missing evidence is load-bearing for the central claim; a published recording or a detailed performance report is required before the claim can be assessed.","section":"Appendix A.1"},{"comment":"The FFT configuration (8192-sample window, 512-sample hop) is stated to have been 'chosen based on experimental results,' but those results are not presented or cited. The 55-dimensional feature vector and this FFT setting are central to the real-time segmentation and matching that supposedly enable the agents to 'dynamically respond' and 'emulate complex musical styles.' Please report the supporting experiments or provide a citation to a fully documented study, and specify the SpireMuse influence-weight settings used in the performance.","section":"Section 2, Offline/Real-time machine listening"},{"comment":"This section is a general literature review on research-creation methodology and AI music and does not provide any evidence about Revival itself. It does not support the paper's central empirical claim, so either the claim must be reduced to a system-description status or this section must be replaced or supplemented with an analysis of the actual performance and interaction data.","section":"Section 3"},{"comment":"The description of the training data is underspecified: the manuscript says the agents were 'trained in works by deceased composers and the collective's compositions,' but it does not identify the corpus, its size, the composers, or how 'training' applies to the self-organizing maps and Markov models of MASOM and SpireMuse. This is necessary to evaluate the claim that the agents 'emulate complex musical styles.'","section":"Section 2 and Appendix A"}],"minor_comments":[{"comment":"The workshop is named 'NeurIPS 2024' in the header but 'NIPS workshop' in Appendix A.1; please use a consistent name.","section":"Header and Appendix A.1"},{"comment":"Reference [8] says 'In the proceedings of the conference on the Institute of Electrical and Electronics Engineers (IEEE)'; this should be 'Proceedings of the IEEE.' Reference [12] appears to have an incomplete title; please verify.","section":"References"},{"comment":"The term 'A/Cs' is ambiguous: it likely means AC power outlets, not air conditioners; please clarify.","section":"Appendix A.2"},{"comment":"Figure 1 (the stage plot) is not referenced in the main text; add a reference or move it closer to the technical description.","section":"Figure 1"},{"comment":"The sentence 'The listening module can adjust feature weights, subsequently influencing the matching algorithms' is vague; it would be helpful to give an example of how the weights are adjusted in real time.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The central issue is evidentiary: the paper claims a realized performance but provides only a proposal-style appendix. This is fixable if the authors can supply a recording, performance log, or independent documentation; alternatively, they could reframe the paper as a system description and remove the unsupported 'demonstrates' claims. The heavy reliance on the authors' own prior systems is acceptable in a workshop/artwork context, but external validation or at least a watchable performance artifact is necessary for an archival venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Reading this as a research-creation workshop paper, what you have is a detailed description of an integrated live audiovisual performance system built from previously published components (MASOM, SpireMuse, Autolume) plus a small-data, ethically motivated framing. The technical setup—machine listening with 55-dimensional features, FFT window/hop choices, OSC communication, Chataigne as conductor, DMX lighting—is concrete and will be useful to practitioners who want to reproduce or adapt this kind of configuration. The advocacy for small, curated datasets and transparency about copyrighted material is a genuine positive stance, and the paper honestly lists the synchronization and autonomy challenges.\n\nThe soft spot is load-bearing: the abstract and conclusion state that Revival 'features real-time co-creative improvisation' and 'demonstrates' potential, but the manuscript gives no recording, no performance log, no audience data, and no analysis of the agents' output. Appendix A.1 says 'we are prepared to perform if needed,' which reads more like a proposal than a report on a realized artwork. So the central empirical claim is asserted, not shown. The FFT configuration is justified only as 'chosen based on experimental results' with no disclosure of those results, so that choice is unverifiable. Section 3 is a generic literature review and adds no direct evidence.\n\nTo be fair, this is a missing-evidence problem, not a sign of fabrication. A short video or time-aligned audio excerpt would go a long way; without it, the paper functions as an artwork statement or project description rather than a demonstration of human-AI co-creation.\n\nWho is this for? Practitioners in the arts and human-AI interaction community who want a component-level description of a live improvisation system. It does not advance the science of musical agents, and the citation pattern is heavily self-referential, which is understandable given the authors are building on their own prior work, but it limits independent verification.\n\nMy recommendation: if the venue explicitly values research-creation practice reports and would accept a major revision that adds documentation, send it out. For a standard CS venue, desk-reject without performance evidence. For the right workshop, it deserves a serious referee, but the referee must be clear that the demonstration claim is currently unsupported.","headline":"A well-described practice report on integrating prior musical agents into a live show, but the paper's central claim that the performance demonstrates co-creation is unsupported by any evidence in the manuscript.","tokens_in":6997,"tokens_out":3359,"would_cite":false,"duration_ms":33660,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI musical agents trained on a small curated corpus can improvise live with human performers, driving music, visuals, and lighting in real time.","keywords":["human-AI co-creation","live audiovisual performance","musical agents","machine listening","small data","self-organizing maps","real-time improvisation","audio-reactive visuals"],"falsifier":"Run the same corpus-matching loop on live audio that lies far outside the training data, such as unpitched percussion-only improvisation or heavily distorted noise, and have independent musicians judge whether the agent's chosen corpus segments are musically related to the input; if those judgments are at chance level, the feature-vector premise of the system fails.","tokens_in":6052,"feed_emoji":"🎶","tokens_out":8442,"duration_ms":85090,"temperature":0.7,"pith_summary":"Revival is a live, roughly thirty-minute audiovisual improvisation in which a percussionist and an electronic musician perform alongside three AI agents, and the paper's thesis is that real-time human-AI co-creation works: the agents, trained on a small curated corpus of the collective's compositions and works by deceased composers, respond to the human players as they play. The paper argues that a machine-listening module that reduces incoming audio to 55-dimensional feature vectors—covering duration, loudness, timbre, pitch, harmony, and emotional valence and arousal—enables the agents to segment the live stream and match it to the corpus in real time. If that works as claimed, the performance stands as a concrete demonstration that small, ethically sourced datasets are enough for generative AI to act as a responsive creative partner in live improvisation. The paper also describes how the same audio features drive reactive visuals and lighting, making the audiovisual layer an output of the same musical interaction rather than a separate technical problem.","feed_headline":"AI agents trained on small datasets improvise live with a human band","feed_subtitle":"Small curated data lets AI agents match live audio in real time, driving music, visuals, and lighting","key_machinery":"The load-bearing mechanism is the machine-listening and corpus-matching loop. Incoming audio is continuously segmented, and each segment is summarized by a 55-dimensional vector comprising duration, mean and standard deviation of loudness, Mel-Frequency Cepstral Coefficients, fundamental frequency, chroma, and valence and arousal from the circumplex model of affect, computed with an FFT window of 8192 samples and a hop size of 512 for high frequency resolution. A self-organizing map organizes the curated corpus, and the real-time matcher aligns incoming segments to the nearest corpus material; one agent, SpireMuse, adds four influence dimensions—rhythmic, spectral, melodic, and harmonic—that let performers weight which features dominate the match. A conductor environment routes OSC messages to the musical agents, the visual synthesizer, and the lighting system, so the same audio feature stream drives both musical response and the audiovisual layer.","core_discovery":"The central claim is that Revival is a working example of real-time co-creative improvisation between human performers and AI musical agents. The agents are built on systems that use self-organizing maps and variable Markov models, or concatenative synthesis with a factor oracle, to organize a curated reference corpus, and a real-time machine-listening loop converts each incoming audio segment into a 55-dimensional feature vector and matches it against that corpus. The paper contends that this lets the agents dynamically respond to human input and emulate complex musical styles, and that the same feature stream, sent over OSC, drives a generative visual synthesizer and DMX lighting under a human VJ's artistic control. The authors present the work as evidence that small-data, ethically curated AI can be a reactive and creative partner in live performance.","pith_inferences":["The paper does not measure whether listeners actually hear the agents as emulating the deceased composers' styles; a listening study comparing perceived style fidelity, musical engagement, and sense of human authorship would test that claim directly.","Because the same feature stream drives agents, visuals, and lighting, the architecture suggests a general pattern for real-time generative art: compress the live input, match it to a curated archive, and re-synthesize across media, a recipe that could transfer to dance, spoken word, or installation art.","The claimed ethical advantage of small data could be examined independently by comparing the resource footprint and copyright clarity of this setup against large-scale generative music systems.","The paper leaves open whether the system would remain coherent with input genres far outside the reference corpus, such as purely percussive improvisation or heavily distorted noise, so the generality of the feature-vector premise remains a testable open question."],"forward_implications":["If the central claim holds, AI musical agents can serve as responsive improvising partners without needing large-scale training corpora.","The performance's roughly thirty-minute structure suggests the interaction loop remains stable over extended co-creation, not just short curated exchanges.","Because visuals and lighting are driven by the same audio features used for musical matching, the audiovisual design is coupled to the musical content by construction.","The four influence dimensions of SpireMuse give performers explicit, real-time control over which musical aspects—rhythm, timbre, melody, or harmony—guide the AI's matching."],"supporting_citations":[{"why":"Supplies the 55-dimensional feature vector, the self-organizing-map clustering, and the core agent architecture used for the percussionist's musical partner.","marker":"[1]"},{"why":"Provides the SpireMuse agent with its real-time machine-listening module and the four influence dimensions (rhythmic, spectral, melodic, harmonic) used in matching.","marker":"[2]"},{"why":"Supplies Autolume, the generative live visual synthesizer that turns audio features into reactive visuals under human VJ control.","marker":"[3]"},{"why":"Establishes the concept of machine listening in interactive music systems that the real-time listening module builds on.","marker":"[5]"},{"why":"Provides the typology and state of the art for musical agents that frames the roles of the three AI agents.","marker":"[6]"},{"why":"Gives the self-organizing-map algorithm that organizes the reference corpus for clustering and matching.","marker":"[8]"},{"why":"Supplies the circumplex model of affect from which the valence and arousal features are computed.","marker":"[9]"},{"why":"Provides the small-data mindset that justifies the curated, ethically sourced corpus and the paper's data-minimalist framing.","marker":"[10]"}],"fun_headline_variants":["AI agents improvise live with human musicians on small data","Small-data AI co-creates music and visuals in real time","Human-AI jam: agents react live to band and drive visuals","AI trained on niche corpora matches live audio to drive show","Live improvisation: AI agents respond to band, feed visuals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The agents' real-time musical responses hinge on the untested premise that a 55-number summary of each short audio segment, computed with fixed spectral-analysis settings, captures enough musical information to match it meaningfully to the curated corpus.","fun_headline_variants_meta":{"raw":{"variants":["AI agents improvise live with human musicians on small data","Small-data AI co-creates music and visuals in real time","Human-AI jam: agents react live to band and drive visuals","AI trained on niche corpora matches live audio to drive show","Live improvisation: AI agents respond to band, feed visuals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000156,"raw_usage":{"total_tokens":1143,"prompt_tokens":798,"completion_tokens":345,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":414,"completion_tokens_details":{"reasoning_tokens":259}},"tokens_in":414,"tokens_out":345,"duration_ms":4240,"temperature":1.0,"reasoning_tokens":259,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:44:54.805768+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same corpus-matching loop on live audio that lies far outside the training data, such as unpitched percussion-only improvisation or heavily distorted noise, and have independent musicians judge whether the agent's chosen corpus segments are musically related to the input; if those judgments are at chance level, the feature-vector premise of the system fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the SpireMuse agent with its real-time machine-listening module and the four influence dimensions (rhythmic, spectral, melodic, harmonic) used in matching."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Autolume, the generative live visual synthesizer that turns audio features into reactive visuals under human VJ control."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the circumplex model of affect from which the valence and arousal features are computed."}],"review_version":1}