{"id":"2e954ffa-75cc-4859-b106-3e1370ec760e","arxiv_id":"2411.13846","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Three trained Hindustani musicians tested a generative vocal model and reported that its output lacked raga, scale, timbre, and style constraints and often felt incoherent with their input.","lead":"This paper reports a pilot user study in which three trained Hindustani musicians tried an AI model that generates vocal melodies. The musicians found the output lacked musical restrictions and did not reliably connect to their input, pointing to constraints and coherence as key design goals.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The two-challenge taxonomy rests on a single-author thematic analysis of three participants, one using an out-of-domain instrument; the categories are not independently verifiable from the paper as written.","rationale":"The paper is carefully framed as an exploratory pilot, and its direct participant quotes support the claim that these three musicians experienced the interaction as lacking constraints and as sometimes incoherent. The reader's CONDITIONAL verdict is therefore appropriate. However, the strongest claim generalizes from three participants to 'trained Hindustani musicians' and 'all three interaction tasks' via a single-author thematic analysis. The reader's weakest_assumption identifies exactly this: three participants and one coder are a thin basis for the two-challenge taxonomy. I agree with that assessment. My additional observation about the timbre sub-claim strengthens the concern: P3's positive reaction to the timbre shift is presented alongside the assertion that participants felt a need for timbre restrictions, which suggests the coding may have imposed a coherent challenge narrative on mixed reactions. A second-coder reliability check would settle whether the taxonomy is stable, and the result would determine whether the claims should remain conditional or be further qualified. Because the reader already recommended CONDITIONAL, my stress-test does not change the verdict.","tokens_in":8124,"tokens_out":6116,"duration_ms":66157,"concrete_test":"Obtain the de-identified interview transcripts from the authors and have a second researcher, blind to the paper's categories, independently perform a Braun and Clarke thematic analysis with a pre-registered codebook containing the candidate themes 'lack of restrictions' and 'incoherence'. Then compute inter-rater agreement (e.g., Cohen's kappa) on which transcript passages instantiate each theme. If kappa is below 0.6, or if the second coder identifies a major additional theme absent from Section 4, the two-challenge taxonomy is not reproducible from the data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that two primary challenges, lack of restrictions and incoherence, characterize trained Hindustani musicians' interaction with GaMaDHaNi. The load-bearing evidence is the thematic analysis described in Sections 2.1 and 2.2: three one-hour interviews coded by a single author, with no codebook, no inter-rater reliability check, and no release of transcripts. The paper then makes universal claims such as 'Throughout all tasks, the participants noted a lack of restrictions' (Section 4.1) and frames these as primary challenges in the abstract, but this rests on three participants whose reactions may be idiosyncratic. The risk is compounded by the study design: one participant (P1) used a harmonium even though GaMaDHaNi was trained only on voice (Section 3.3), so findings about scale, timbre, and coherence from that participant may reflect an out-of-distribution input modality rather than a general property of the interaction. There is also a notable internal tension in the timbre sub-claim: Section 4.1 says participants felt a need for timbre restrictions, yet P3 explicitly described the female-sounding timbre shift as 'delightful'. That contradiction suggests the thematic coding may have merged divergent reactions into a single challenge category. Since the paper presents no second coder, no agreement metric, and no larger sample, the two-challenge taxonomy cannot be distinguished from the interpretive preferences of one analyst on a very small, partly out-of-distribution dataset.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative pilot study in which three trained Hindustani musicians interacted with GaMaDHaNi, a hierarchical generative model for Hindustani vocal contours, through three predefined interaction tasks: idea generation, call and response, and melodic reinterpretation. The authors identify two primary challenges from participant reactions and discussions: (1) a lack of restrictions in the model output (with respect to raga, scale, timbre, and style) and (2) incoherence between user input and model output. The paper situates these challenges in the context of Hindustani music, discusses creativity and constraints in that tradition, and suggests future design directions such as conditioning mechanisms. The study is explicitly framed as an exploratory, in-the-wild pilot, acknowledging that the model was not adapted for the interaction setting and that one participant used an out-of-distribution harmonium input.","tokens_in":8384,"tokens_out":4046,"duration_ms":39719,"significance":"If the reported findings are taken as an exploratory hypothesis-generating study, the paper makes a useful contribution by documenting concrete expectations and frustrations of trained Hindustani musicians with a generative model, an underexplored area in human-AI interaction. The direct participant quotes and the authors' explicit acknowledgment of the out-of-distribution nature of the study are strengths, as is the grounding in Hindustani music theory (raga, scale, style, creativity under constraints). The proposed future directions—such as conditioning on raga and scale—are reasonable and likely to inform subsequent work. However, the significance is tempered by the very small sample size (n=3) and the reliance on a single-author thematic analysis with no reported codebook or inter-rater reliability check, which limits the confidence with which the two-challenge taxonomy can be treated as a robust finding rather than an interpretive reading of a few individuals' experiences.","major_comments":[{"comment":"","section":"Section 2.2 and Section 4"},{"comment":"","section":"Section 3.3 and Section 4.1"},{"comment":"","section":"Section 4.1, Timbre paragraph"}],"minor_comments":[{"comment":"","section":"Section 4.2"},{"comment":"","section":"Section 4.2 and 5"},{"comment":"","section":"Figure 2 caption"},{"comment":"","section":"References"},{"comment":"","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an exploratory qualitative study with a very small sample size, but it is honestly framed as a pilot. The main concern is that the paper's core taxonomy is derived from a single-author thematic analysis without transparency regarding the coding process. If the authors are willing to soften the strength of their claims, provide more methodological detail, and clearly separate the out-of-distribution participant's data, the paper could be suitable for publication as a pilot study in a creative-AI or human-AI interaction venue. I also note that the authors are evaluating their own model; while this is not inherently problematic, the single-coder analysis exacerbates the risk of interpretive bias, and the report should encourage the authors to address this explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a pilot study, honestly framed as one, and the two headline challenges - lack of restrictions and incoherence - are real observations from the material, not a fabrication. The authors put three trained Hindustani musicians in front of GaMaDHaNi, report what they said, and use pitch-contour examples to illustrate. For a domain with almost no prior human-AI interaction work, that is a legitimate first step.\n\nThe genuinely new thing: no prior work I know of reports practicing musicians' reactions to a continuous-pitch generative model for Hindustani music. The findings that musicians expect raga/scale adherence and coherent call-and-response, and the distinction between low-level (scale) and higher-level (mood, musical idea) attributes, are useful for anyone building conditioning or evaluation for such models. The model is the authors' own prior work, but evaluating it is not circular; the challenges come from participant quotes, not assumptions in the model.\n\nThe soft spots are real but not fatal. The thematic analysis is a single author, no codebook, no second coder, no released transcripts. With n = 3, the two-challenge taxonomy is not independently verifiable. I don't think the central claim collapses - the direct quotes carry the argument - but the paper would be stronger with a clearer coding procedure or longer excerpts. The harmonium participant (P1) is an out-of-distribution input, and the authors flag it themselves. One nuance: the timbre sub-claim is more mixed than the text implies. Not all participants wanted timbre restriction; P3 found the female-sounding timbre 'delightful'. The paper acknowledges different reactions, but the challenge category is glossed a bit too uniformly.\n\nWho this is for: people working on human-AI co-creation in music, especially for non-Western traditions, and anyone building conditioning for continuous-pitch generative models. I'd send it to review - it deserves referee time - but I'd ask the authors to soften the generality, describe the coding process in more detail, and consider releasing anonymized transcripts.","headline":"A clean, honest pilot study: three trained musicians interact with a continuous-pitch Hindustani music model, and the two reported challenges are real but the evidence base is a single-coder thematic analysis of a very small sample.","tokens_in":8907,"tokens_out":2550,"would_cite":true,"duration_ms":26357,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Trained Hindustani musicians who interacted with the generative model GaMaDHaNi reported that its output lacked raga, scale, timbre, and style constraints and often did not coherently respond to their musical input.","keywords":["Hindustani music","generative music model","human-AI interaction","raga constraints","pitch contours","user study","diffusion models","call and response"],"falsifier":"Run a larger interaction study in which each of, say, twenty trained Hindustani musicians uses the same three tasks, have two independent analysts code the transcripts, and separately measure how often the model's output stays in the input's scale and raga; if most participants find scale and raga maintained in most outputs, or if independent coders do not reproduce the two-challenge structure, the claim is overturned.","tokens_in":7913,"feed_emoji":"🎵","tokens_out":6005,"duration_ms":50812,"temperature":0.7,"pith_summary":"This paper is a pilot study of how practicing Hindustani musicians interact with GaMaDHaNi, a hierarchical generative model that produces vocal pitch contours and audio from continuous fundamental-frequency contours. Three trained musicians, one vocalist and two harmonium players, tried three interactions: choosing continuations, call-and-response, and melodic reinterpretation. The paper's central finding is that participants experienced the model's output as insufficiently constrained, wandering across ragas and scales, shifting timbre, and sometimes crossing into a different musical style. They also found the output incoherent as a response, failing to continue the scale, raga, mood, or musical idea they offered. The paper reads these as challenges for model design and proposes raga-, scale-, timbre-, and style-aware conditioning as future directions.","feed_headline":"Musicians: AI raga model needs more rules and consistency","feed_subtitle":"A pilot study finds a raga model's free-flowing output ignores scale, timbre, and style, and often fails to respond.","key_machinery":"The object carrying the study is GaMaDHaNi, a two-level hierarchical generative model whose intermediate representation is a continuous fundamental-frequency pitch contour. A Pitch Generator produces the contour; a Spectrogram Generator conditioned on the contour and singer identity produces a mel-spectrogram; Griffin-Lim converts it to audio. Interaction is implemented through primed generation, which treats the last $t_{\\text{prime}}$ seconds of a user input as the start of a fixed 12-second generation, and through melodic reinterpretation, which uses Iterative $\\alpha$-Deblending reverse diffusion from a user pitch-contour guide to synthesize a new contour. The diffusion formulation $x_{\\alpha} = (1-\\alpha)x_0 + \\alpha x_1$ and the reverse step are the mechanism by which user input enters generation, and the paper's argument is that this mechanism, without added musical constraints, produces the observed lack of restrictions and incoherence.","core_discovery":"In its own terms, the study establishes that an unadapted generative model for Hindustani vocal music, presented to expert practitioners, is judged less as a creative partner than as a companion that violates the tradition's implicit rules. Across all three tasks participants identified two primary difficulties: the model lacked restrictions on raga, scale, timbre, and style, and its output was inconsistent with their input, failing to maintain scale, structure, mood, or musical idea. The paper argues that this points to a deeper need: in Hindustani music creativity is idiomatic, operating within fixed seed ideas such as raga, taal, and performance frameworks, so a generative tool must impose constraints to be usable. It therefore recommends conditioning and a robust discrete-note mapping as the path toward coherent, tradition-respecting interaction.","pith_inferences":["I would infer that the two challenges are likely to persist for any model that conditions only on pitch-contour primes, since nothing in that conditioning mechanism carries raga, scale, timbre, style, or response-coherence information.","A quantitative diagnostic suggested by the paper's examples would be to measure scale adherence and raga-resemblance statistics over many generated responses and see whether the complaints concentrate in the unconstrained generation tasks.","The same 'lack of restrictions' could be reframed as a tunable creative parameter, with tight raga adherence for practice and looser generation for exploration, rather than simply a defect.","A minimal coherence test follows from P2's description of a good response: a model that repeats the input phrase with slight variation would likely satisfy much of the call-and-response expectation and is easy to prototype."],"forward_implications":["A usable interactive Hindustani music generator will need to generate under constraints derived from raga, scale, timbre, and style, not just from a raw pitch-contour prime.","Coherent call-and-response requires the model to preserve both low-level attributes such as scale and rhythmic structure and high-level attributes such as mood and the repetition of a musical idea.","Because the model works on continuous pitch contours, imposing note-level constraints first requires a robust mapping between discrete notes and ornamented, microtonal continuous contours.","Future deployments should treat interaction input as a distinct distribution from the ensemble recording data used in training and plan for that shift explicitly."],"supporting_citations":[{"why":"Supplies GaMaDHaNi, the hierarchical generative model whose interactive behavior the study evaluates.","marker":"Shikarpur et al. [2024]"},{"why":"Supplies the thematic analysis procedure used to extract the two primary challenges from participant transcripts.","marker":"Braun and Clarke [2006]"},{"why":"Supplies Iterative Alpha-Deblending, the diffusion objective used by GaMaDHaNi's generators and by the melodic reinterpretation mechanism.","marker":"Heitz et al. [2023]"},{"why":"Supplies the SDEdit guided-generation idea adapted for the melodic reinterpretation task.","marker":"Meng et al. [2022]"},{"why":"One of the two datasets the model was trained on, defining the training distribution the study treats as out of distribution.","marker":"Srinivasamurthy et al. [2021]"},{"why":"The other training dataset, also defining the distribution gap between training and interaction input.","marker":"Gulati et al. [2016]"}],"fun_headline_variants":["AI raga model needs rules: study finds free output confuses musicians","Raga AI lacks constraints: pilot sees musicians want structure","Hindustani music AI: untamed output fails expert ears","Generative raga model must follow rules to satisfy musicians","Pilot: AI raga vocalists need conditioning for coherence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study's central claim rests on three musicians' reactions being representative and on a single researcher's thematic analysis being accurate; if either fails, the two challenges would not generalize.","fun_headline_variants_meta":{"raw":{"variants":["AI raga model needs rules: study finds free output confuses musicians","Raga AI lacks constraints: pilot sees musicians want structure","Hindustani music AI: untamed output fails expert ears","Generative raga model must follow rules to satisfy musicians","Pilot: AI raga vocalists need conditioning for coherence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":1164,"prompt_tokens":849,"completion_tokens":315,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":465,"completion_tokens_details":{"reasoning_tokens":229}},"tokens_in":465,"tokens_out":315,"duration_ms":3607,"temperature":1.0,"reasoning_tokens":229,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:48:30.512330+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a larger interaction study in which each of, say, twenty trained Hindustani musicians uses the same three tasks, have two independent analysts code the transcripts, and separately measure how often the model's output stays in the input's scale and raga; if most participants find scale and raga maintained in most outputs, or if independent coders do not reproduce the two-challenge structure, the claim is overturned.","supporting_citations":[],"review_version":1}