{"id":"72f6c803-351b-4e37-974d-66a7d342dccc","arxiv_id":"2412.06469","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A proof-of-concept embodied AI system using a trained neural-network factory and a robotic arm that draws or gestures to co-improvise with a mixed ensemble, with qualitative reports of transformed and inclusive music-making.","lead":"Jess+ is an AI-driven robotic arm system that acts as a digital score for a mixed ensemble of disabled and non-disabled musicians, sensing audio, EEG, and skin conductance to generate drawing and movement gestures in live improvisation. A four-month case study with three musicians suggests the system supported inclusive co-creative experiences, though the authors say they cannot yet explain the causal mechanism.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Train/serve skew in the AI Factory: deployment normalises each buffer separately and substitutes robot-tip coordinates for human core, so the seven models' outputs are not valid predictors; the embodied-AI coupling claim is unsupported.","rationale":"The reader correctly identified the AI Factory's prediction quality as the weakest assumption. My stress test confirms this but locates a more concrete and internal problem: even before considering cross-context transfer from eight pianists to the trio, the deployment pipeline is inconsistent with the training pipeline in two ways that invalidate the model outputs. The per-buffer normalisation is explicitly stated in the Deployment section, so this is not speculation; it is a train/serve skew. The robot-tip-as-core substitution is also explicit in the same section and is not accompanied by any calibration. These issues mean the 'thought train' streams are not meaningful estimates of flow, core position, or audio envelope, so the gesture manager's behaviour is largely decoupled from the musicians' state. This does not destroy the paper's value as a design case study: the musicians reported rich, positive experiences, and the authors themselves include a strong uncertainty caveat. However, it does undermine the stronger causal claim in the conclusion. The conditional verdict remains appropriate, and the proposed offline test is feasible because the dataset and code are open-source. No issue with the qualitative methods themselves is raised here; the concern is about the technical validity of the AI component that the causal claim relies upon.","tokens_in":13162,"tokens_out":4184,"duration_ms":46788,"concrete_test":"Offline validation using the open-source dataset and code: take the held-out validation chunks from the Embodied Musicking Dataset and run the seven models under the exact deployment preprocessing (per-buffer min-max normalisation of each input feature, no global training-set scaling) and compare each predicted stream to the held-out ground-truth target (e.g., Spearman correlation and RMSE against a mean-predictor baseline). In parallel, test the robot-tip substitution by re-estimating model 3's flow prediction using robot-workspace coordinates under the same transform used live and checking whether correlation with self-reported flow remains above chance. If per-buffer normalisation reduces correlations to near zero, the AI Factory carries no usable signal in deployment, so the gesture manager's choices cannot be attributed to the embodied-AI models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal claim is that the implemented embodied-AI design decisions 'led to rich experiences... transformed their practice.' That attribution requires the AI Factory outputs to be meaningful predictors when deployed. Two deployment mismatches break this before any generalisation question arises. First, the models were trained with min-max scaling computed per-channel across the whole training set, but Deployment explicitly normalises 'across the 5-sec buffer (as opposed to the whole training set).' This train/serve skew forces every buffer to span [0,1], destroying absolute amplitude information for the audio envelope and EDA, and can produce out-of-distribution inputs whenever live buffer extrema differ from the training extrema. The predicted flow/core streams are therefore not the quantities the models were trained to estimate. Second, models 2 and 3 were trained on the human core mid-shoulder x,y position, but deployment feeds the robot arm tip Cartesian position as a 'robotic representation of the core mid-shoulder position.' No coordinate alignment or calibration is reported; live robot coordinates live in a different physical space, so those predictions are not about the musicians' bodies. Since the gesture manager selects one of the seven model streams (plus audio and random streams) and applies 0.1/0.7 thresholds to drive gestures, its behaviour is effectively random or audio-threshold-driven. The qualitative experience may still be positive, and the authors' caveat that they 'cannot say with any certainty' is appropriate; but the conclusion that the embodied-AI approach 'did contribute' overreaches the evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents Jess+, an intelligent digital score system that uses a robotic arm as an embodied AI co-creative partner for a mixed-ability ensemble. The system combines audio, EEG, and EDA sensing with an 'AI Factory' of seven neural networks trained on an 'Embodied Musicking Dataset' of eight pianists improvising to a jazz backing track. The networks predict flow, body position, and audio features from one another; a gesture manager selects among the resulting 'thought-train' streams and a belief system maps them to robotic gestures. The authors report qualitative findings from four iterative workshops and a final sharing performance with three musicians (including one disabled musician), and they claim in the abstract and conclusion that the implemented design decisions and embodied-AI approach led to rich experiences that transformed the musicians' practice. The paper also states, in the Discussion, that the authors 'can not say with any certainty how they led to such transformations encounters for the musicians' and that a deep dive into dataset correlations reveals 'extremely loose causality.' The paper is positioned as a companion to a CHI user-experience paper and aims to provide the technical details necessary for reproducibility.","tokens_in":13428,"tokens_out":6443,"duration_ms":62657,"significance":"If the central causal claim were supported, Jess+ would be a notable contribution to inclusive music technology and human-robot co-creativity, integrating physiological sensing, multiple neural networks, and a robotic arm in a modular, open-source system, and reporting positive experiences from a real mixed-ability ensemble. The paper's strengths include its clear architectural description, the inclusion of an open-source implementation, and its honest acknowledgment of uncertainty in parts of the Discussion. However, the significance is currently undercut by the lack of validation of the deployed ML pipeline: the training and deployment normalization procedures differ, and the robot arm position is used as a proxy for human body position without calibration. Because the gesture manager mixes model outputs with random and audio-threshold streams, the robot's behavior may be effectively decoupled from the AI predictions, weakening the attribution of the musicians' experiences to the embodied-AI design. The qualitative testimonials support a modest claim about rich user experience, but they do not, as presented, support the causal claim made in the abstract and conclusion.","major_comments":[{"comment":"The deployed models are not fed the same input distribution that the models were trained on. The Training section states that 'Data was normalised with min-max feature scaling by computing the minimums and maximums for each channel of each feature across the training set,' while the Deployment section states that features 'are normalised in real-time across the 5-sec buffer (as opposed to the whole training set).' This train/serve skew means that every live buffer is rescaled to [0,1] independently, destroying absolute amplitude information for the audio envelope and EDA, and can produce out-of-distribution inputs whenever the live buffer extrema differ from the training extrema. The model outputs used to drive the robot are therefore not the quantities the models were trained to estimate. The paper needs to provide evidence that the deployed pipeline produces meaningful predictions (for example, by evaluating deployment normalization on held-out data or comparing predictions against random baselines), or the causal claim that the AI Factory's output drove the musicians' experience must be withdrawn.","section":"Deployment (near Figure 4)"},{"comment":"Models 2, 3, and 6 were trained on the x and y positions of the human core mid-shoulder point, but in deployment model 3 is fed 'the x and y positions of the robot arm tip' as a 'robotic representation of the core mid-shoulder position.' No coordinate alignment, calibration, or reference-frame mapping is reported; the robot tip is in a different physical space (pen on a table or drawing board, arm mounted on a pallet) from the musician's shoulders. Feeding this substitution into model 3, and using the predicted core as a target for model 6, yields predictions that are not about the musicians' bodies. This is a load-bearing gap for the 'self-awareness' stream (items c and f in Figure 4), and it must be either validated or removed from the causal account.","section":"Deployment, models 3 and 6 (Figure 4)"},{"comment":"The paper validates the seven models only through training/validation loss curves (Figure 3 shows MSE loss for two models) and manual hyperparameter selection. There is no evaluation on held-out data in terms of prediction quality (e.g., correlation, R-squared, classification accuracy), no comparison to trivial baselines, and no per-model analysis. Since the gesture manager selects among the seven model streams plus audio and random streams (Gesture manager section, with the audio stream given 36% probability and the other streams equally probable), and applies 0.1/0.7 thresholds, it is plausible that the robot's behavior is dominated by the audio-threshold 'startle' and random choice rather than by the neural-network predictions. Without task-relevant validation of the deployed models, the paper's assertion that the AI Factory 'seems to be key' (Discussion) is unsupported.","section":"Training and Figure 3"},{"comment":"The causal attribution in the abstract and conclusion ('the implemented design decisions and embodied-AI approach led to rich experiences... transformed their practice') is not supported by the study design. The three musicians co-designed the belief system and gesture language during four iterative workshops, so their positive testimonials may reflect ownership, novelty, and the collaborative development process rather than the specific embodied-AI mechanisms. There is no control condition (e.g., a robot with scripted or random gestures) and no comparison to other digital-score systems. The authors themselves acknowledge in the Discussion that 'we can not say with any certainty how they led to such transformations encounters for the musicians' and that the dataset correlations show 'extremely loose causality.' The abstract and conclusion should be aligned with this more modest evidence; alternatively, the paper should present the qualitative findings as a design case study rather than as confirmation of the causal claim.","section":"User centred design and Results"},{"comment":"The AI Factory models are trained on eight pianists improvising to a jazz backing track (Dataset), but are deployed with a trio of musicians (Ableton, violin, cello) engaged in free improvisation with no backing track (Results). The modal and stylistic shift is large, and the paper provides no analysis of whether the learned correlations between body movement, physiological response, and audio envelope transfer to this context. This generalizability gap compounds the train/serve skew identified above, and it should be addressed explicitly, for example by reporting how the predicted streams behave on live data or by adding a domain-adaptation discussion.","section":"Dataset and Deployment"}],"minor_comments":[{"comment":"The phrase 'holds this stream for a few sections' should read 'holds this stream for a few seconds'; the intended time unit is clear from the later 'gesture phrases of 3 to 8 seconds.'","section":"Layer 3 - gesture manager"},{"comment":"In the Deployment list, 'predicted flow from d) if fed into model 7)' should be 'predicted flow from d) is fed into model 7)'; the same typo appears in the Features/Models list where model 7 is described.","section":"Deployment, item (g)"},{"comment":"The musician's name is spelled 'Clare' in the Results section but 'Claire' in the Discussion when quoting her; please make the spelling consistent.","section":"Results and Discussion"},{"comment":"The sentence 'we can not say with any certainty how they led to such transformations encounters for the musicians' is grammatically incomplete; 'transformations encounters' appears to be a typo for 'transformational encounters.'","section":"Introduction and Conclusion"},{"comment":"Several cross-references are empty: 'described in )' and 'discussed in )' in the 'Interaction loop' section, and the reference to Figure 2 appears as a placeholder. Please fill in the figure and section numbers before submission.","section":"Interaction loop and modular design"},{"comment":"In 'Every 0.5 to 2 seconds (selected randomly) which represent a rhythmic loop,' the verb 'represent' should agree with the singular subject 'which'; use 'represents.'","section":"Gesture manager"}],"recommendation":"major_revision","confidential_remarks":"The paper is a system/design contribution with a companion CHI paper covering the user-experience analysis. The technical architecture and open-source software are valuable, but the current framing overclaims causal evidence, and the AI Factory deployment suffers from train/serve and proxy-input issues that are not merely cosmetic. I recommend major revision rather than rejection, because the system itself is a worthwhile contribution and the problems could be addressed by (a) validating the deployed ML pipeline or explicitly acknowledging its limitations, (b) reframing the central claim as a design case study rather than a causal finding, and (c) aligning the abstract and conclusion with the authors' own caveats in the Discussion. If the authors are unwilling to soften the causal claim or provide additional validation, rejection would be appropriate; with those revisions, the paper could make a legitimate contribution to the HCI and inclusive music technology literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Craig, here's my read of Vear & Benerradi's Jess+ paper.\n\nWhat's new and worth taking seriously: the integration itself. A robotic arm as a digital score, driven by seven parallel encoder-decoder models trained on an 'embodied musicking' dataset, plus EEG/EDA from the disabled musician, plus a belief-system gesture manager — that combination doesn't appear in the cited prior art. The paper gives enough architectural detail to reproduce the system, and the code is open source under BSD-2. That's real, formal, reproducible work, and I'd cite it if I were working on accessible music technology or HRI for co-creativity.\n\nThe qualitative material is also handled carefully. The authors quote the musicians, distinguish between the disabled and non-disabled musicians' relationships with the robot, and repeatedly hedge that they can't say with certainty how the design led to the reported transformations. That honesty is to their credit, and the testimonials are plausible as descriptions of rich experiences.\n\nWhere it gets soft — and this is where I think the stress-test note is right, not wrong: the deployment of the AI Factory is not faithful to the training setup. The models were trained on min-max scaling computed across the whole training set; at deployment they normalize each 5-second buffer independently. That's a textbook train/serve skew, and it means the audio envelope and EDA inputs at deployment are not the same quantities the models learned to predict from. On top of that, models 2 and 3 were trained on human core mid-shoulder coordinates, but at deployment they're fed robot-arm tip positions with no reported coordinate alignment or calibration. Those predictions aren't about the musicians' bodies. Since the gesture manager tosses a coin between nine streams anyway (audio at 36%, the rest equal), the robot's behavior is largely random or audio-threshold-driven. The claim that the embodied-AI approach 'did contribute' to the transformations therefore overreaches: the system as a whole contributed, but the specific coupling via the neural models isn't demonstrated.\n\nIs this fatal? Not for the paper as a design report. But it is load-bearing for the conclusion. The fix is straightforward in principle: an offline validation of the seven models on held-out data, or a clear reframing that treats the AI Factory as an aesthetic component whose predictive validity wasn't assessed. The co-design circularity — musicians helped shape the belief system and gesture language, then their positive feedback is cited as evidence — is a minor additional concern, not the main problem.\n\nWho's this for? Researchers in accessible music tech and human-robot co-creativity. It deserves a serious referee — the system is novel and the reporting is honest — but the referee should demand either the validation or a humbler conclusion.","headline":"A genuinely novel integrated system and a careful design report, but the central causal claim outruns the evidence, and the deployment pipeline has a train/serve skew the authors don't address.","tokens_in":13990,"tokens_out":1801,"would_cite":true,"duration_ms":18402,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A robotic arm guided by embodied AI served as a digital score that let a mixed-ability ensemble improvise together as co-creators.","keywords":["embodied AI","inclusive music-making","digital score","human-robot interaction","brain-computer music interface","musicking","robotic arm","co-creative improvisation"],"falsifier":"Play the same four improvisation pieces with the gesture manager fed genuine live sensor streams versus the same streams shuffled or replaced by random numbers while keeping the gesture library and thresholds identical; if the musicians cannot distinguish the two conditions, or their reports of co-creation do not change, the central claim of meaningful embodied-AI coupling is refuted.","tokens_in":12915,"feed_emoji":"🤖","tokens_out":5744,"duration_ms":54419,"temperature":0.7,"pith_summary":"This paper argues that an embodied-AI system called Jess+, in which a robotic arm is driven by neural networks trained on recordings of pianists improvising, can function as an intelligent digital score for a mixed ensemble of disabled and non-disabled musicians. The claim is that the design decisions—a modular AI Factory, a gesture manager, and an embedded belief system—led to deeply engaged experiences for the musicians and transformed their practice as an inclusive ensemble. The musicians reported being in-the-loop and perceiving the robot as a co-creative partner; the disabled musician described it as an extension of herself, while the others saw it as a creative accompanist. The authors are careful to say they cannot explain with certainty how the system produced these transformations, only that embedding music-specific embodied interaction data into each design layer contributed to them.","feed_headline":"Robotic arm co-creates music with disabled and non-disabled players","feed_subtitle":"Seven neural nets on live audio, EEG and skin signals drive the gestures that inspired the ensemble's improvisation.","key_machinery":"The load-bearing mechanism is the closed-loop interaction design with a four-layer modular architecture. Layer 1 formats audio, EEG, EDA, and robot-arm positions; Layer 2, the AI Factory, runs seven hourglass-shaped convolutional neural networks that predict one feature stream from another, such as audio-to-flow, EEG-to-flow, and flow-to-core; Layer 3, the gesture manager, selects among nine streams (seven model outputs, live audio amplitude, and a random poetry stream) and maps their values to low, medium, or high responses; Layer 4, the belief system, converts those values into a predefined gestural language of shapes and graphic-score-inspired movements with randomly varied speed, size, and other performance parameters. The gesture manager also has a startled response that interrupts a gesture phrase when live sound exceeds a threshold, creating a two-way interaction where the robot both dances to the musicians and conducts them.","core_discovery":"The central discovery is that a robot arm can serve as a co-creative digital score in live improvisation when its behaviour is coupled to a closed loop of sound, physiological signals, and a library of gestures. Jess+ senses the ensemble's audio plus EEG and electrodermal activity from the disabled musician, feeds these through seven small neural networks trained on an Embodied Musicking Dataset of eight pianists, and converts their outputs into thought trains that select movements of a pen- or feather-wielding robotic arm. The musicians reported back-and-forth interaction in real time, and a public sharing performance demonstrated the system working with a live audience. The authors assert that embedding music-specific embodied interaction data and behaviours into every design layer contributed to the musicians' transformational encounters, though the exact causal pathway remains unknown.","pith_inferences":["Beyond the paper's claims, the system's design suggests a general template for embodied co-creative agents: a perception layer connected to a factory of simple predictors, a stochastic selector, and a curated expressive language, a template that could transfer to other art forms such as dance or theatre.","The paper leaves open whether the neural networks' predictions are the active ingredient. A direct test would be to run the same workshops with the seven model streams replaced by random values while keeping the gesture library and thresholds; if the musicians' experience is unchanged, the 'AI' contribution would be shown to be decorative.","The startled response and the 0.1/0.7 thresholds introduce a turn-taking and interruption mechanism that could be studied as a form of human-robot coordination, potentially informing non-musical assistive and collaborative robotics.","The Embodied Musicking Dataset, collected from only eight pianists improvising to a jazz backing track, is a narrow basis for a system used in free improvisation; collecting a more diverse dataset and validating model predictions on the deployment context would be a natural next step the paper does not take."],"forward_implications":["If the embodied-AI claim holds, disabled musicians can participate in live improvisation as full co-creators rather than being limited by the interface barriers of traditional instruments.","A robotic arm using physiological and audio sensing can act as a non-judgemental 'third space' that reduces the psychological pressure of human-to-human improvisation, as the musicians reported.","The modular, subsumption-inspired architecture allows individual components—sensors, models, gesture library—to be replaced or updated without rebuilding the whole system, supporting iterative co-design with musicians.","The open-source release of Jess+ makes the system reproducible for other inclusive music-making projects.","The closed-loop design, including the startled response, offers a concrete model for how an embodied agent can alternate between following and leading in a collaborative improvisation."],"supporting_citations":[{"why":"Companion study whose qualitative analysis of the musicians' reflections supplies the experiential findings this paper interprets.","marker":"Vear et al. 2024"},{"why":"Defines 'musicking' as taking part, grounding the project's aim of meaning-making through participation.","marker":"Small 1998"},{"why":"Provides the subsumption architecture that the modular AI stack is modeled on.","marker":"Brooks 1991"},{"why":"Supplies the definition of embodied AI used throughout the design.","marker":"Vear 2022"},{"why":"Frames the 'belief system' as an acceptance that something is true, justifying the embedded aesthetic of the digital score.","marker":"Barr, Feigenbaum, and Cohen 1981"},{"why":"Motivates the 'trains of thought' metaphor behind running seven separate networks rather than a single multi-variable model.","marker":"Gelernter 2010"}],"fun_headline_variants":["Robot arm conducts inclusive jam using EEG and trained neural nets","Jess+ robot arm turns brainwaves and audio into live music gestures","Embodied AI robot arm co-creates with mixed ensemble in real time","Neural nets and EEG inspire robot arm in inclusive music improvisation","Inclusive music ensemble thrives with AI robot arm co-composer"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that the seven neural networks, trained on eight pianists improvising to a jazz backing track, still produce meaningful real-time predictions when fed the live audio, EEG, EDA, and robot-arm positions of a different trio improvising freely; if those predictions are meaningless, the robot's behaviour reduces to loudness-triggered random gesture selection.","fun_headline_variants_meta":{"raw":{"variants":["Robot arm conducts inclusive jam using EEG and trained neural nets","Jess+ robot arm turns brainwaves and audio into live music gestures","Embodied AI robot arm co-creates with mixed ensemble in real time","Neural nets and EEG inspire robot arm in inclusive music improvisation","Inclusive music ensemble thrives with AI robot arm co-composer"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000417,"raw_usage":{"total_tokens":2096,"prompt_tokens":839,"completion_tokens":1257,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":455,"completion_tokens_details":{"reasoning_tokens":1168}},"tokens_in":455,"tokens_out":1257,"duration_ms":12065,"temperature":1.0,"reasoning_tokens":1168,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:36:20.857394+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Play the same four improvisation pieces with the gesture manager fed genuine live sensor streams versus the same streams shuffled or replaced by random numbers while keeping the gesture library and thresholds identical; if the musicians cannot distinguish the two conditions, or their reports of co-creation do not change, the central claim of meaningful embodied-AI coupling is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Companion study whose qualitative analysis of the musicians' reflections supplies the experiential findings this paper interprets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines 'musicking' as taking part, grounding the project's aim of meaning-making through participation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the subsumption architecture that the modular AI stack is modeled on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the definition of embodied AI used throughout the design."},{"cited_title":"A.; and Cohen, P","cited_arxiv_id":null,"evidence_quote":"Frames the 'belief system' as an acceptance that something is true, justifying the embedded aesthetic of the digital score."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the 'trains of thought' metaphor behind running seven separate networks rather than a single multi-variable model."}],"review_version":1}