{"id":"93fd7812-9457-4ebb-bb36-7c5eef8820ff","arxiv_id":"2502.00023","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper presents MACAT and MACataRT, enhanced corpus-based musical agents that add real-time synthesis, visualization, and factor-oracle temporal modeling to the existing MASOM and CataRT systems.","lead":"This paper describes two AI music systems, MACAT and MACataRT, that improvise alongside human musicians by recombining small, personalized audio corpora. It is a system and practice report from a research lab, with live performances as evidence rather than quantitative evaluation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing baseline comparison leaves the 'broaden creative options' claim unsupported despite credible existence evidence.","rationale":"The reader's weakest-assumption analysis and my own converge: the paper's evidence establishes that the systems were built and used, but not that they deliver the claimed creative broadening over existing baselines. I read the paper in good faith and credit the concrete artifacts: the GitHub repository link, detailed interface descriptions, live performance citations, and the AI Music Song Contest finalist result are real supporting evidence for existence. My concern is not that the authors are dishonest or that the systems are fake; it is that the paper's final claim is an empirical effectiveness claim that outruns the evidence offered. I also noticed a secondary internal inconsistency between Section 2.1, which says MACAT learns the sequence of nodes using a VMM, and Appendix B.2, which says the current architecture is not equipped with the VMM but uses the factor oracle. I did not make that the headline concern because it affects the precision of the architecture description rather than the central functional claim, and it could be a typo. The reader's CONDITIONAL verdict is exactly right: the systems are plausibly real, but accepting the full conclusion requires the type of comparative evidence that is currently absent. My recommended outcome is UNCHANGED, because the conditionality already captures this gap.","tokens_in":9884,"tokens_out":6316,"duration_ms":69565,"concrete_test":"Conduct a preregistered, counterbalanced studio study with at least 12 musician participants: each participant improvises for equal time with MACAT versus MASOM, and MACataRT versus CataRT, using the same personal audio corpus. Collect preference ratings, counts of excerpts judged musically usable, Likert ratings of perceived co-creativity, and open-ended justifications. If MACAT and MACataRT are not preferred at a pre-specified threshold or do not yield more usable material than their baselines, the 'broaden creative options' conclusion fails. A minimal alternative is to analyze the YouTube performance and contest entry with a rubric assessing whether factor-oracle-driven sections produce output that MASOM or CataRT could not plausibly produce on the same corpus.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion that MACAT and MACataRT 'significantly broaden the creative options for musicians' is never tested against the two systems they extend: MASOM and CataRT. The paper documents the systems' components, interfaces, and live uses, including a contest finalist, but those demonstrations only show that the systems can run in performance. They do not show that adding the factor oracle or real-time synthesis changes improvisational outcomes for the better, and they certainly do not support 'significantly.' The authors explicitly decline quantitative evaluation in Section 3.1, but they also provide no user study, listening test, structured rehearsal observation, or analysis of the cited performance recordings relative to baseline behavior. Because 'functional systems exist' and 'these systems broaden creative options' are distinct claims, the latter remains unverified. This is not an internal inconsistency, but it is the assumption on which the paper's value claim depends. If the paper claimed only that two working tools are described, the absence of evaluation would be acceptable; as written, the conclusion overreaches the evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents two MAX/MSP musical agent systems, MACAT and MACataRT, which extend the authors' prior MASOM system and IRCAM's CataRT with real-time concatenative synthesis and factor-oracle-based temporal modeling. It describes the systems' workflows, interfaces, machine listening components, and an ethical rationale built around small, personalized audio corpora. The paper reports live performances, including a Top-10 finalist entry in the 2024 AI Music Song Contest, and argues that these systems 'significantly broaden the creative options for musicians' in real-time co-creative improvisation.","tokens_in":10040,"tokens_out":4780,"duration_ms":44872,"significance":"The main strength is the detailed, well-illustrated description of two apparently functional open-source systems, with publicly available code and real-world performance documentation. The design philosophy of small data and artist-in-the-loop control is timely and relevant to the ethical AI-in-music discourse. However, the paper's stated value claim—that the systems 'significantly broaden creative options'—is not supported by any user study, listening test, or baseline comparison; as it stands the contribution is a credible systems/design report rather than an empirically validated claim about creative expansion.","major_comments":[{"comment":"The sentence 'MACAT and MACataRT demonstrate how artist-in-the-loop AI agents can significantly broaden the creative options for musicians' is a load-bearing claim that the evidence does not currently support. The paper presents no user study, listening test, structured rehearsal observation, or analysis of the cited performance recordings. Section 3.1 explicitly declines quantitative evaluation, but the conclusion makes a comparative and magnitude claim ('significantly broaden') that requires some form of evidence. Either temper the claim to a demonstration of functional tools or add an evaluation component such as expert assessment, performance transcripts, or a comparison of MACataRT with CataRT on a defined musical task.","section":"Section 4 / Abstract"},{"comment":"The claimed advantage of MACataRT over CataRT is the addition of a factor oracle temporal model to address CataRT's lack of temporal structure. This is described and shown in workflow diagrams, but no evidence shows that the factor oracle produces improved temporal coherence or creative utility relative to CataRT's KNN-based selection. Add at least a small case study or baseline comparison—for example, generation from the same corpus with and without the factor oracle—to make this load-bearing design claim credible.","section":"Section 2.2 / Figure 1(d)"},{"comment":"The ethical claims—that small personalized datasets make training data 'straightforward tracking' and 'foster trust and confidence'—are presented as conclusions, but the paper provides no user data or systematic traceability analysis. As a design rationale these statements are acceptable, but as assertions about user experience and transparency they require either evidence or explicit reframing as intended properties.","section":"Appendix A"}],"minor_comments":[{"comment":"The phrase 'Echoes of Synthetic F orest' contains a typo; it should read 'Forest'.","section":"Section 3.2"},{"comment":"The title 'Adaptive Concatenative Sound Synthesis and Its Application to Micromontage Compositior' has a typo; 'Compositior' should be 'Composition'.","section":"Reference [7]"},{"comment":"The acronym 'MFFC' should be 'MFCC' in the first bullet under extracted audio features.","section":"Appendix B.3"},{"comment":"Please clarify the relationship between VMM and the factor oracle in generation: Section 2.1 says MACAT learns sequences with VMM and uses FO for pattern recognition, while Appendix B.2 says the current MASOM architecture 'is not equipped with the VOMM ... but with Factor Oracle.' This appears inconsistent and should be reconciled.","section":"Section 2.1 vs. Appendix B.2"},{"comment":"The text refers to 'Equation 2a' when discussing spectral flatness statistics, but the equations are numbered (1) and (2); the reference should be corrected.","section":"Appendix B.3"}],"recommendation":"major_revision","confidential_remarks":"This is a workshop-style systems paper. The central engineering contribution is credible and the code is publicly available, but the framing overclaims: the 'significantly broaden creative options' conclusion needs either revision or minimal evaluative evidence. I recommend major revision rather than rejection, as the gap can be fixed by narrowing the conclusion or adding a modest comparison study. The authors should also clarify the novelty relative to their previous MASOM and CataRT work, since the paper currently leaves this to workflow diagrams."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick read of arXiv:2502.00023. The genuinely new thing is two working MAX/MSP extensions: MACAT grafts real-time concatenative synthesis and a visual interface onto the MASOM agent, and MACataRT adds a factor-oracle temporal model to CataRT. The paper gives a detailed, specific walkthrough of interfaces, parameters, and the two operational modes (reactive and proactive). It also reports real uses, including a Top-10 finalist in the 2024 AI Music Song Contest. That is real existence proof: these are not vaporware, and the GitHub repo makes them accessible.\n\nWhat it does well: the system descriptions in the appendices are concrete enough to reproduce the tools, and the small-data training argument is coherent. The ethical transparency discussion is mostly assertion, but it is grounded in a real design choice (personalized corpora) rather than hand-waving.\n\nSoft spots: the central conclusion claims these systems \"significantly broaden the creative options\" for musicians. The paper's own methodology explicitly disclaims quantitative evaluation, and there is no user study, listening test, or structured observation. The performances show the systems ran, not that they beat MASOM or CataRT. The stress-test note is right: the factor-oracle addition is plausible but never compared against a no-oracle baseline. So \"significantly\" is unsupported. That is a real gap, but it is a scope issue, not bad faith. The paper would be more honest if it presented itself as a system/practice report rather than evidence of improved creative outcomes.\n\nWho gets value: researchers in musical metacreation and interactive performance, especially anyone building on CataRT or MASOM. This is workshop-level, not a field-reorganizing result. I would still give it a serious referee because the systems are real, clearly described, and the practice documentation is useful. My recommendation: send it out, but ask the authors to either add a baseline comparison or tone down the \"significantly broadens\" claim.\n\nBest.","headline":"A useful system/practice report on two working MAX/MSP musical agents, but the conclusion overreaches: the factor-oracle and real-time synthesis additions are plausible, not demonstrated as better than baseline.","tokens_in":10539,"tokens_out":1529,"would_cite":true,"duration_ms":16596,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MACAT and MACataRT are real-time musical agents that improvise with humans using small personal audio corpora.","keywords":["musical agents","concatenative synthesis","factor oracle","self-organizing map","audio mosaicing","real-time improvisation","small data","co-creative AI"],"falsifier":"A controlled listening study comparing MACataRT's proactive mode against the same system with the factor oracle disabled and the segment order randomized would settle whether the temporal model adds musical coherence; if listeners cannot distinguish the oracle version from the scrambled baseline, the central creative claim is unsupported.","tokens_in":9693,"feed_emoji":"🎵","tokens_out":9542,"duration_ms":84894,"temperature":0.7,"pith_summary":"The paper presents two working musical agents, MACAT and MACataRT, built for real-time human–AI co-creation. The central claim is that both systems let a musician bring a small corpus of their own recordings and get an improvising partner that responds live, either leading a solo piece (MACAT) or collaborating through reactive and proactive audio mosaicing (MACataRT). The authors argue that this small-data, artist-in-the-loop approach preserves expressive nuance, keeps training transparent and ethical, and broadens what a performer can do in live improvisation. They offer live performances, including a top-ten finalist entry in the 2024 AI Music Song Contest, as evidence that the systems function in real musical practice.","feed_headline":"Live musical agents learn from small personal audio sets","feed_subtitle":"Concatenative synthesis plus factor-oracle memory lets a machine improvise from a musician's own recordings.","key_machinery":"The load-bearing mechanism is the pairing of concatenative sound synthesis with the factor oracle. Concatenative synthesis rebuilds sound by selecting and stitching short segments from a user-supplied corpus according to audio descriptors, preserving the expressive character of the original recordings. The factor oracle is a suffix automaton, trained offline on the sequence of segment (or SOM node) indices, that recognizes recurring patterns and generates new sequences by forward or backward jumps; a probability parameter controls how often the oracle moves forward rather than backwards. MACAT adds a self-organizing map—a 2D grid that clusters perceptually similar segments—and a machine-listening loop that feeds the current audio state back into the oracle. MACataRT instead applies the oracle to the indices produced by CataRT-style audio mosaicing, so the same corpus can drive either reactive accompaniment or autonomous continuation.","core_discovery":"The paper's central claim is that MACAT and MACataRT are functioning musical agent systems for real-time co-creative improvisation, built on corpus-based concatenative synthesis and trained on small, personalized audio corpora. MACAT is the agent-led system: it clusters corpus segments with a self-organizing map, learns the node sequence with a variable-order Markov model, and uses a factor oracle plus a self-listening loop to choose what to play next, with a congruence parameter that lets the performer choose between repetition and exploration. MACataRT is the collaborative system: it keeps CataRT-style audio mosaicing and descriptor targeting, and adds a factor oracle trained on the sequence of segment indices, giving it two modes, reactive improvisation (respond to the live input) and proactive improvisation (continue from learned patterns). The paper offers live performances, including a top-ten finalist piece in the 2024 AI Music Song Contest, as evidence that the systems work in musical practice.","pith_inferences":["Beyond the paper: the reactive/proactive split suggests a general design pattern—the same corpus and synthesizer can serve both as an accompanist and as an autonomous continuation engine, and musicians' preferences between those two roles are a testable question the paper does not isolate.","Beyond the paper: since every generated segment is drawn from a known corpus, the architecture could log segment indices and produce an exact audit trail of which source recordings contributed to a performance, making the claimed transparency mechanically verifiable.","Beyond the paper: a natural next comparison is to pit these small-data agents against a large-scale generative model on stylistic coherence with the same performer; the paper's design implies the small-data systems would win on personalization and traceability."],"forward_implications":["A musician with a small set of personal recordings can train an improvising agent on a laptop CPU, with no GPU or large dataset, and use it in live performance.","MACataRT's factor-oracle layer turns a reactive mosaicing instrument into one that can also generate autonomously, so the same corpus supports both following the human and leading the music.","MACAT's self-listening feedback allows a single performer to play alongside a machine that hears, clusters, and re-sequences its own output in real time.","Because all source material comes from a small curated corpus, a performance's audio provenance remains traceable, which supports the paper's ethical transparency argument."],"supporting_citations":[{"why":"It supplies the CataRT corpus-based concatenative synthesis engine and audio descriptor targeting that MACataRT extends.","marker":"[6]"},{"why":"It supplies the MASOM architecture—SOM memory, machine listening, and sequence learning—from which MACAT is derived.","marker":"[14]"},{"why":"It provides the factor oracle suffix automaton used by both agents for real-time pattern recognition and generation.","marker":"[17]"},{"why":"It defines feature-driven audio mosaicing, the reassembly technique underlying MACataRT's reactive mode.","marker":"[19]"},{"why":"It motivates the small-data mindset behind training on personalized, curated corpora.","marker":"[10]"},{"why":"It supplies the working definition of musical agents and the task space the paper situates itself in.","marker":"[1]"},{"why":"It provides the variable-order Markov model that MACAT uses alongside the factor oracle for sequence recognition.","marker":"[16]"},{"why":"It provides the self-organizing map that clusters audio segments into the nodes MACAT visualizes and sequences.","marker":"[15]"}],"fun_headline_variants":["MACAT and MACataRT: two AI agents improvise from your audio","Real-time co-creative AI music from a musician's own small corpus","Agent-led vs collaborative: two musical AI systems from personal sets","Small personal audio drives live AI improvisation in two ways","AI music agents use factor oracle and self-listening from your tracks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that these systems broaden creative options rests on treating successful live performances as sufficient evidence, because the paper deliberately sets aside quantitative evaluation.","fun_headline_variants_meta":{"raw":{"variants":["MACAT and MACataRT: two AI agents improvise from your audio","Real-time co-creative AI music from a musician's own small corpus","Agent-led vs collaborative: two musical AI systems from personal sets","Small personal audio drives live AI improvisation in two ways","AI music agents use factor oracle and self-listening from your tracks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000304,"raw_usage":{"total_tokens":1704,"prompt_tokens":863,"completion_tokens":841,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":750}},"tokens_in":479,"tokens_out":841,"duration_ms":7911,"temperature":1.0,"reasoning_tokens":750,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:32:27.526872+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled listening study comparing MACataRT's proactive mode against the same system with the factor oracle disabled and the segment order randomized would settle whether the temporal model adds musical coherence; if listeners cannot distinguish the oracle version from the scrambled baseline, the central creative claim is unsupported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the CataRT corpus-based concatenative synthesis engine and audio descriptor targeting that MACataRT extends."},{"cited_title":"& Pasquier, P","cited_arxiv_id":null,"evidence_quote":"It supplies the MASOM architecture—SOM memory, machine listening, and sequence learning—from which MACAT is derived."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the factor oracle suffix automaton used by both agents for real-time pattern recognition and generation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It defines feature-driven audio mosaicing, the reassembly technique underlying MACataRT's reactive mode."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It motivates the small-data mindset behind training on personalized, curated corpora."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the working definition of musical agents and the task space the paper situates itself in."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the variable-order Markov model that MACAT uses alongside the factor oracle for sequence recognition."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the self-organizing map that clusters audio segments into the nodes MACAT visualizes and sequences."}],"review_version":1}