REVIEW 3 major objections 5 minor 35 references
Exploratory Study Of Human-AI Interaction For Hindustani Music
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Trained Hindustani musicians who interacted with the generative model GaMaDHaNi reported that its output lacked raga, scale, timbre, and style constraints and often did not coherently respond to their musical input.
desk verdict A clean, honest pilot study: three trained musicians interact with a continuous-pitch Hindustani music model, and the two reported challenges are real but the evidence base is a single-coder thematic analysis of a very small sample. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The object carrying the study is GaMaDHaNi, a two-level hierarchical generative model whose intermediate representation is a continuous fundamental-frequency pitch contour. A Pitch Generator produces the contour; a Spectrogram Generator conditioned on the contour and singer identity produces a mel-spectrogram; Griffin-Lim converts it to audio. Interaction is implemented through primed generation, which treats the last $t_{\text{prime}}$ seconds of a user input as the start of a fixed 12-second generation, and through melodic reinterpretation, which uses Iterative $\alpha$-Deblending reverse diffusion from a user pitch-contour guide to synthesize a new contour. The diffusion formulation $x_{\alpha} = (1-\alpha)x_0 + \alpha x_1$ and the reverse step are the mechanism by which user input enters generation, and the paper's argument is that this mechanism, without added musical constraints, produces the observed lack of restrictions and incoherence.
What would settle it
Run a larger interaction study in which each of, say, twenty trained Hindustani musicians uses the same three tasks, have two independent analysts code the transcripts, and separately measure how often the model's output stays in the input's scale and raga; if most participants find scale and raga maintained in most outputs, or if independent coders do not reproduce the two-challenge structure, the claim is overturned.
Extended reading notes
Core claim
In its own terms, the study establishes that an unadapted generative model for Hindustani vocal music, presented to expert practitioners, is judged less as a creative partner than as a companion that violates the tradition's implicit rules. Across all three tasks participants identified two primary difficulties: the model lacked restrictions on raga, scale, timbre, and style, and its output was inconsistent with their input, failing to maintain scale, structure, mood, or musical idea. The paper argues that this points to a deeper need: in Hindustani music creativity is idiomatic, operating within fixed seed ideas such as raga, taal, and performance frameworks, so a generative tool must impose constraints to be usable. It therefore recommends conditioning and a robust discrete-note mapping as the path toward coherent, tradition-respecting interaction.
Load-bearing premise
The study's central claim rests on three musicians' reactions being representative and on a single researcher's thematic analysis being accurate; if either fails, the two challenges would not generalize.
Editorial extensions
If this is right
- A usable interactive Hindustani music generator will need to generate under constraints derived from raga, scale, timbre, and style, not just from a raw pitch-contour prime.
- Coherent call-and-response requires the model to preserve both low-level attributes such as scale and rhythmic structure and high-level attributes such as mood and the repetition of a musical idea.
- Because the model works on continuous pitch contours, imposing note-level constraints first requires a robust mapping between discrete notes and ornamented, microtonal continuous contours.
- Future deployments should treat interaction input as a distinct distribution from the ensemble recording data used in training and plan for that shift explicitly.
Reading between the lines
- I would infer that the two challenges are likely to persist for any model that conditions only on pitch-contour primes, since nothing in that conditioning mechanism carries raga, scale, timbre, style, or response-coherence information.
- A quantitative diagnostic suggested by the paper's examples would be to measure scale adherence and raga-resemblance statistics over many generated responses and see whether the complaints concentrate in the unconstrained generation tasks.
- The same 'lack of restrictions' could be reframed as a tunable creative parameter, with tight raga adherence for practice and looser generation for exploration, rather than simply a defect.
- A minimal coherence test follows from P2's description of a good response: a model that repeats the input phrase with slight variation would likely satisfy much of the call-and-response expectation and is easy to prototype.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a qualitative pilot study in which three trained Hindustani musicians interacted with GaMaDHaNi, a hierarchical generative model for Hindustani vocal contours, through three predefined interaction tasks: idea generation, call and response, and melodic reinterpretation. The authors identify two primary challenges from participant reactions and discussions: (1) a lack of restrictions in the model output (with respect to raga, scale, timbre, and style) and (2) incoherence between user input and model output. The paper situates these challenges in the context of Hindustani music, discusses creativity and constraints in that tradition, and suggests future design directions such as conditioning mechanisms. The study is explicitly framed as an exploratory, in-the-wild pilot, acknowledging that the model was not adapted for the interaction setting and that one participant used an out-of-distribution harmonium input.
Significance. If the reported findings are taken as an exploratory hypothesis-generating study, the paper makes a useful contribution by documenting concrete expectations and frustrations of trained Hindustani musicians with a generative model, an underexplored area in human-AI interaction. The direct participant quotes and the authors' explicit acknowledgment of the out-of-distribution nature of the study are strengths, as is the grounding in Hindustani music theory (raga, scale, style, creativity under constraints). The proposed future directions—such as conditioning on raga and scale—are reasonable and likely to inform subsequent work. However, the significance is tempered by the very small sample size (n=3) and the reliance on a single-author thematic analysis with no reported codebook or inter-rater reliability check, which limits the confidence with which the two-challenge taxonomy can be treated as a robust finding rather than an interpretive reading of a few individuals' experiences.
major comments (3)
- [Section 2.2 and Section 4]
- [Section 3.3 and Section 4.1]
- [Section 4.1, Timbre paragraph]
minor comments (5)
- [Section 4.2]
- [Section 4.2 and 5]
- [Figure 2 caption]
- [References]
- [Section 3.2]
Circularity Check
No significant circularity: the two identified challenges are empirical findings from a small qualitative user study, not derivations from model inputs or self-cited results.
full rationale
The paper's central claims are the two participant-identified challenges: lack of restrictions and inconsistency/incoherence of model output. These are presented as outcomes of a thematic analysis of interview transcripts (Section 2.2: 'A thematic analysis (Braun and Clarke [2006]) was performed on the interview transcripts'), supported by direct participant quotations in Sections 4.1 and 4.2. The claims do not reduce by construction to any input parameter, fitting procedure, or equation. The diffusion equations in Section 3.2 describe the model's generation mechanism but are not used to derive the challenges; they are background for how the tasks were implemented. The paper does cite the authors' own prior model, GaMaDHaNi (Shikarpur et al. [2024]), but that citation identifies the system under study rather than supplying the paper's findings. No uniqueness theorem or load-bearing self-citation forces the two-challenge taxonomy. The paper explicitly acknowledges the study's limitations, including the out-of-distribution nature of the interaction and one participant using a harmonium despite voice-only training (Section 3.3), which further confirms that the findings are treated as empirical observations rather than as analytically forced outcomes. Concerns about the small sample, single-coder thematic analysis, or generalizability are validity or evidentiary concerns, not circularity. There is no fitted input renamed as a prediction, no quantity defined in terms of the quantity it is said to explain, and no known result merely relabeled. The derivation chain, such as it is, runs from participant reactions to the challenge categories, with no circular step.
Assumptions & free parameters
free parameters (2)
- t0 (diffusion start step) =
0.5
- Task timing parameters =
tprime=2s (Task 1), tprime=4s (Task 2), output truncation=4s, response=8s
assumptions (3)
- domain assumption Thematic analysis (Braun and Clarke, 2006) is an appropriate method for extracting challenges from interview transcripts.
- domain assumption The model outputs encountered in the study are representative of GaMaDHaNi's intended behavior.
- domain assumption Raga, scale, timbre, and style are the salient dimensions along which Hindustani musicians evaluate melodic output.
Cite this review
Pith. "Pith review of Exploratory Study Of Human-AI Interaction For Hindustani Music." pith.science (2026). https://pith.science/paper/HNPYHKL2
@misc{pith2026241113846,
author = {Pith},
title = {Pith review of: Exploratory Study Of Human-AI Interaction For Hindustani Music},
year = {2026},
howpublished = {\url{https://pith.science/paper/HNPYHKL2}},
note = {Machine review of arXiv:2411.13846}
}
read the original abstract
This paper presents a study of participants interacting with and using GaMaDHaNi, a novel hierarchical generative model for Hindustani vocal contours. To explore possible use cases in human-AI interaction, we conducted a user study with three participants, each engaging with the model through three predefined interaction modes. Although this study was conducted "in the wild"- with the model unadapted for the shift from the training data to real-world interaction - we use it as a pilot to better understand the expectations, reactions, and preferences of practicing musicians when engaging with such a model. We note their challenges as (1) the lack of restrictions in model output, and (2) the incoherence of model output. We situate these challenges in the context of Hindustani music and aim to suggest future directions for the model design to address these gaps.
Figures
Reference graph
Works this paper leans on
-
[1]
A. Abid, A. Abdalla, A. Abid, D. Khan, A. Alfozan, and J. Zou. Gradio: Hassle-free sharing and testing of ml models in the wild. arXiv preprint arXiv:1906.02569, 2019
arXiv 1906
-
[2]
S. Adhikary, M. S. M, S. S. K, S. Bhat, and K. P. L. Automatic music generation of indian classical music based on raga. In Proc. of the IEEE International Conference for Convergence in Technology (I2CT), 2023
work page 2023
-
[3]
V. Braun and V. Clarke. Using thematic analysis in psychology. Qualitative research in psychology, 2006
work page 2006
-
[4]
T. B. Brown, B. Mann, N. Ryder, and M. Subbaiah. Language models are few-shot learners. arXiv preprint ArXiv:2005.14165, 2020
arXiv 2005
- [5]
- [6]
-
[7]
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341, 2020
arXiv 2005
- [8]
Show all 35 references
-
[9]
Engel, M
J. Engel, M. Hoffman, and A. Roberts. Latent constraints: Learning to generate conditionally from unconditional generative models. arXiv preprint arXiv:1711.05772, 2017
2017 arXiv
-
[10]
F. Font, A. P, G. Stegmann, and F. Brinkmann. URL https://github.com/AudioCommons/timbral_models?tab=readme-ov-file. Accessed on 2024-8-16
2024
-
[11]
K. K. Ganguli. Guruji padmashree pt. ajoy chakrabarty: As i have seen him. Samakalika Sangeetham, 2012
2012
-
[12]
Gopi and F
S. Gopi and F. William. Introductory studies on raga multi-track music generation of indian classical music using ai. The International Conference on AI and Musical Creativity, 2023
2023
-
[13]
Griffin and J
D. Griffin and J. Lim. Signal estimation from modified short-time fourier transform. IEEE Transactions on acoustics, speech, and signal processing, 32 0 (2): 0 236--243, 1984
1984
-
[14]
Gulati, J
S. Gulati, J. Serr \`a Juli \`a , K. K. Ganguli, S. Sent \"u rk, and X. Serra. Time-delayed melody surfaces for r \=a ga recognition. In Devaney J, Mandel MI, Turnbull D, Tzanetakis G, editors. ISMIR 2016. Proceedings of the 17th International Society for Music Information Ret...
2016
-
[15]
Haught-Tromp
C. Haught-Tromp. The green eggs and ham hypothesis: How constraints facilitate creativity. Psychology of Aesthetics, Creativity, and the Arts, 2017
2017
-
[16]
Heitz, L
E. Heitz, L. Belcour, and T. Chambon. Iterative -(de) blending: A minimalist deterministic diffusion model. In Proc. of the ACM SIGGRAPH Conference, 2023
2023
-
[17]
Ho and T
J. Ho and T. Salimans. Classifier-free diffusion guidance. In NeurIPS Workshop on Deep Generative Models and Downstream Applications, 2021
2021
-
[18]
C.-Z. A. Huang, T. Cooijmans, A. Roberts, A. Courville, and D. Eck. Counterpoint by convolution. In International Society for Music Information Retrieval (ISMIR), 2019
2019
-
[19]
D. Huron. Sweet anticipation: Music and the psychology of expectation. MIT press, 2008
2008
-
[20]
N. A. Jairazbhoy. The r \=a gs of North Indian music: their structure and evolution . Popular Prakashan, 1995
1995
-
[21]
Jaques, S
N. Jaques, S. Gu, R. E. Turner, and D. Eck. Tuning recurrent neural networks with reinforcement learning. In International Conference on Learning Representations, 2017
2017
-
[22]
L. Leante. Observing musicians/audience interaction in north indian classical music performance. In Musicians and their Audiences. Routledge, 2016
2016
-
[23]
A. McNeil. Seed ideas and creativity in hindustani raga music: beyond the composition--improvisation dialectic. In Ethnomusicology Forum. Taylor & Francis, 2017
2017
-
[24]
C. Meng, Y. He, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon. SDE dit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations, 2022
2022
-
[25]
Nikrang, K
A. Nikrang, K. Breckner, T. Neumayr, F. Hirschmann, and M. Augstein. Ai creativity in the light of autonomy. Artificial Intelligence, Co-Creation and Creativity: The New Frontier for Innovation, 2024
2024
-
[26]
H. Patel. The Importance of the Teacher , 11 2007. URL https://rhythmicthoughts.wordpress.com/tag/guru-mukhi-vidya/. Accessed on 2024-08-16
2007
-
[27]
H. S. Powers and R. Widdess. Theory and practice of classical music. In India, subcontinent of. Grove Music, 2001
2001
-
[28]
H. V. Sahasrabuddhe. Analysis and synthesis of hindustani classical music, 1992. URL https://www.cse.iitb.ac.in/ hvs/paper_1992.html. Accessed: 2024-07-22
1992
-
[29]
Shikarpur, K
N. Shikarpur, K. M. Dendukuri, Y. Wu, A. Caillon, and C.-Z. A. Huang. Hierarchical generative modeling of melodic vocal contours in hindustani classical music. In International Society for Music Information Retrieval (ISMIR), 2024
2024
-
[30]
Srinivasamurthy, S
A. Srinivasamurthy, S. Gulati, R. C. Repetto, and X. Serra. Saraga: open datasets for research on indian art music. Empirical Musicology Review, 16 0 (1): 0 85--98, 2021
2021
-
[31]
Vidwans, P
A. Vidwans, P. Verma, and P. Rao. Classifying cultural music using melodic features. In International Conference on Signal Processing and Communications (SPCOM). IEEE, 2012
2012
-
[32]
V. Vidwans. Computational music, 2017. URL https://computationalmusic.com/. Accessed on 2024-04-11
2017
-
[33]
R. Widdess. Language, Music, and Interaction, chapter Schemas and improvisation in Indian music. College Publications, 2013
2013
-
[34]
S.-L. Wu, C. Donahue, S. Watanabe, and N. J. Bryan. Music controlnet: Multiple time-varying controls for music generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024
2024
-
[35]
Zhang, A
L. Zhang, A. Rao, and M. Agrawala. Adding conditional control to text-to-image diffusion models. In IEEE/CVF International Conference on Computer Vision, 2023
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.