Pith. sign in

REVIEW 3 major objections 5 minor 35 references

Exploratory Study Of Human-AI Interaction For Hindustani Music

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Trained Hindustani musicians who interacted with the generative model GaMaDHaNi reported that its output lacked raga, scale, timbre, and style constraints and often did not coherently respond to their musical input.

desk verdict A clean, honest pilot study: three trained musicians interact with a continuous-pitch Hindustani music model, and the two reported challenges are real but the evidence base is a single-coder thematic analysis of a very small sample. read the letter →

arxiv 2411.13846 v1 pith:HNPYHKL2 submitted 2024-11-21 cs.HC cs.AI

classification cs.HCcs.AI
keywords Hindustanimusicgenerativemodelhuman-AIinteractionragaconstraintspitchcontoursuserstudydiffusionmodelscallandresponse
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a pilot study of how practicing Hindustani musicians interact with GaMaDHaNi, a hierarchical generative model that produces vocal pitch contours and audio from continuous fundamental-frequency contours. Three trained musicians, one vocalist and two harmonium players, tried three interactions: choosing continuations, call-and-response, and melodic reinterpretation. The paper's central finding is that participants experienced the model's output as insufficiently constrained, wandering across ragas and scales, shifting timbre, and sometimes crossing into a different musical style. They also found the output incoherent as a response, failing to continue the scale, raga, mood, or musical idea they offered. The paper reads these as challenges for model design and proposes raga-, scale-, timbre-, and style-aware conditioning as future directions.

What carries the argument

The object carrying the study is GaMaDHaNi, a two-level hierarchical generative model whose intermediate representation is a continuous fundamental-frequency pitch contour. A Pitch Generator produces the contour; a Spectrogram Generator conditioned on the contour and singer identity produces a mel-spectrogram; Griffin-Lim converts it to audio. Interaction is implemented through primed generation, which treats the last $t_{\text{prime}}$ seconds of a user input as the start of a fixed 12-second generation, and through melodic reinterpretation, which uses Iterative $\alpha$-Deblending reverse diffusion from a user pitch-contour guide to synthesize a new contour. The diffusion formulation $x_{\alpha} = (1-\alpha)x_0 + \alpha x_1$ and the reverse step are the mechanism by which user input enters generation, and the paper's argument is that this mechanism, without added musical constraints, produces the observed lack of restrictions and incoherence.

What would settle it

Run a larger interaction study in which each of, say, twenty trained Hindustani musicians uses the same three tasks, have two independent analysts code the transcripts, and separately measure how often the model's output stays in the input's scale and raga; if most participants find scale and raga maintained in most outputs, or if independent coders do not reproduce the two-challenge structure, the claim is overturned.

Watch

Extended reading notes

Core claim

In its own terms, the study establishes that an unadapted generative model for Hindustani vocal music, presented to expert practitioners, is judged less as a creative partner than as a companion that violates the tradition's implicit rules. Across all three tasks participants identified two primary difficulties: the model lacked restrictions on raga, scale, timbre, and style, and its output was inconsistent with their input, failing to maintain scale, structure, mood, or musical idea. The paper argues that this points to a deeper need: in Hindustani music creativity is idiomatic, operating within fixed seed ideas such as raga, taal, and performance frameworks, so a generative tool must impose constraints to be usable. It therefore recommends conditioning and a robust discrete-note mapping as the path toward coherent, tradition-respecting interaction.

Load-bearing premise

The study's central claim rests on three musicians' reactions being representative and on a single researcher's thematic analysis being accurate; if either fails, the two challenges would not generalize.

Editorial extensions

If this is right

  • A usable interactive Hindustani music generator will need to generate under constraints derived from raga, scale, timbre, and style, not just from a raw pitch-contour prime.
  • Coherent call-and-response requires the model to preserve both low-level attributes such as scale and rhythmic structure and high-level attributes such as mood and the repetition of a musical idea.
  • Because the model works on continuous pitch contours, imposing note-level constraints first requires a robust mapping between discrete notes and ornamented, microtonal continuous contours.
  • Future deployments should treat interaction input as a distinct distribution from the ensemble recording data used in training and plan for that shift explicitly.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I would infer that the two challenges are likely to persist for any model that conditions only on pitch-contour primes, since nothing in that conditioning mechanism carries raga, scale, timbre, style, or response-coherence information.
  • A quantitative diagnostic suggested by the paper's examples would be to measure scale adherence and raga-resemblance statistics over many generated responses and see whether the complaints concentrate in the unconstrained generation tasks.
  • The same 'lack of restrictions' could be reframed as a tunable creative parameter, with tight raga adherence for practice and looser generation for exploration, rather than simply a defect.
  • A minimal coherence test follows from P2's description of a good response: a model that repeats the input phrase with slight variation would likely satisfy much of the call-and-response expectation and is easy to prototype.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a qualitative pilot study in which three trained Hindustani musicians interacted with GaMaDHaNi, a hierarchical generative model for Hindustani vocal contours, through three predefined interaction tasks: idea generation, call and response, and melodic reinterpretation. The authors identify two primary challenges from participant reactions and discussions: (1) a lack of restrictions in the model output (with respect to raga, scale, timbre, and style) and (2) incoherence between user input and model output. The paper situates these challenges in the context of Hindustani music, discusses creativity and constraints in that tradition, and suggests future design directions such as conditioning mechanisms. The study is explicitly framed as an exploratory, in-the-wild pilot, acknowledging that the model was not adapted for the interaction setting and that one participant used an out-of-distribution harmonium input.

Significance. If the reported findings are taken as an exploratory hypothesis-generating study, the paper makes a useful contribution by documenting concrete expectations and frustrations of trained Hindustani musicians with a generative model, an underexplored area in human-AI interaction. The direct participant quotes and the authors' explicit acknowledgment of the out-of-distribution nature of the study are strengths, as is the grounding in Hindustani music theory (raga, scale, style, creativity under constraints). The proposed future directions—such as conditioning on raga and scale—are reasonable and likely to inform subsequent work. However, the significance is tempered by the very small sample size (n=3) and the reliance on a single-author thematic analysis with no reported codebook or inter-rater reliability check, which limits the confidence with which the two-challenge taxonomy can be treated as a robust finding rather than an interpretive reading of a few individuals' experiences.

major comments (3)
  1. [Section 2.2 and Section 4]
  2. [Section 3.3 and Section 4.1]
  3. [Section 4.1, Timbre paragraph]
minor comments (5)
  1. [Section 4.2]
  2. [Section 4.2 and 5]
  3. [Figure 2 caption]
  4. [References]
  5. [Section 3.2]

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the two identified challenges are empirical findings from a small qualitative user study, not derivations from model inputs or self-cited results.

full rationale

The paper's central claims are the two participant-identified challenges: lack of restrictions and inconsistency/incoherence of model output. These are presented as outcomes of a thematic analysis of interview transcripts (Section 2.2: 'A thematic analysis (Braun and Clarke [2006]) was performed on the interview transcripts'), supported by direct participant quotations in Sections 4.1 and 4.2. The claims do not reduce by construction to any input parameter, fitting procedure, or equation. The diffusion equations in Section 3.2 describe the model's generation mechanism but are not used to derive the challenges; they are background for how the tasks were implemented. The paper does cite the authors' own prior model, GaMaDHaNi (Shikarpur et al. [2024]), but that citation identifies the system under study rather than supplying the paper's findings. No uniqueness theorem or load-bearing self-citation forces the two-challenge taxonomy. The paper explicitly acknowledges the study's limitations, including the out-of-distribution nature of the interaction and one participant using a harmonium despite voice-only training (Section 3.3), which further confirms that the findings are treated as empirical observations rather than as analytically forced outcomes. Concerns about the small sample, single-coder thematic analysis, or generalizability are validity or evidentiary concerns, not circularity. There is no fitted input renamed as a prediction, no quantity defined in terms of the quantity it is said to explain, and no known result merely relabeled. The derivation chain, such as it is, runs from participant reactions to the challenge categories, with no circular step.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central claim rests on a small qualitative study. No fitted model parameters are introduced in this paper, but the study's task timing and diffusion start time are hand-set. The main axioms are methodological: thematic analysis validity, representativeness of the model, and the salience of raga, scale, timbre, and style dimensions.

free parameters (2)
  • t0 (diffusion start step) = 0.5
    Equation 2 fixes the reverse diffusion start at t0=0.5 for the melodic reinterpretation task; this hand-set value controls how strongly the output follows the user guide and shapes the task experience.
  • Task timing parameters = tprime=2s (Task 1), tprime=4s (Task 2), output truncation=4s, response=8s
    Section 3.1 sets these durations by hand. They define the interaction modes and could influence perceived coherence, though the central qualitative claim does not depend on the exact values.
assumptions (3)
  • domain assumption Thematic analysis (Braun and Clarke, 2006) is an appropriate method for extracting challenges from interview transcripts.
    Invoked in Section 2.2 as the analysis method; the paper does not justify why this qualitative method is valid for this context.
  • domain assumption The model outputs encountered in the study are representative of GaMaDHaNi's intended behavior.
    Section 3 treats the Gradio-based interface and the three tasks as a faithful instantiation of the model; interaction bugs or implementation artifacts are not discussed.
  • domain assumption Raga, scale, timbre, and style are the salient dimensions along which Hindustani musicians evaluate melodic output.
    Section 4.1 organizes participant challenges under these categories, but the paper provides no independent evidence that these dimensions are exhaustive or the most important.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exploratory Study Of Human-AI Interaction For Hindustani Music." pith.science (2026). https://pith.science/paper/HNPYHKL2

@misc{pith2026241113846,
  author       = {Pith},
  title        = {Pith review of: Exploratory Study Of Human-AI Interaction For Hindustani Music},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HNPYHKL2}},
  note         = {Machine review of arXiv:2411.13846}
}
read the original abstract

This paper presents a study of participants interacting with and using GaMaDHaNi, a novel hierarchical generative model for Hindustani vocal contours. To explore possible use cases in human-AI interaction, we conducted a user study with three participants, each engaging with the model through three predefined interaction modes. Although this study was conducted "in the wild"- with the model unadapted for the shift from the training data to real-world interaction - we use it as a pilot to better understand the expectations, reactions, and preferences of practicing musicians when engaging with such a model. We note their challenges as (1) the lack of restrictions in model output, and (2) the incoherence of model output. We situate these challenges in the context of Hindustani music and aim to suggest future directions for the model design to address these gaps.

Figures

Figures reproduced from arXiv: 2411.13846 by the authors.

Figure 1
Figure 1. The hierarchical structure of the generative model, [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Extracted pitch contours of call and response examples where coherence was maintained [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

35 extracted references · 31 canonical work pages

  1. [1]

    A. Abid, A. Abdalla, A. Abid, D. Khan, A. Alfozan, and J. Zou. Gradio: Hassle-free sharing and testing of ml models in the wild. arXiv preprint arXiv:1906.02569, 2019

  2. [2]

    Adhikary, M

    S. Adhikary, M. S. M, S. S. K, S. Bhat, and K. P. L. Automatic music generation of indian classical music based on raga. In Proc. of the IEEE International Conference for Convergence in Technology (I2CT), 2023

  3. [3]

    Braun and V

    V. Braun and V. Clarke. Using thematic analysis in psychology. Qualitative research in psychology, 2006

  4. [4]

    T. B. Brown, B. Mann, N. Ryder, and M. Subbaiah. Language models are few-shot learners. arXiv preprint ArXiv:2005.14165, 2020

  5. [5]

    Das and M

    D. Das and M. Choudhury. Finite state models for generation of hindustani classical music. In International Symposium on Frontiers of Research in Speech and Music, 2005

  6. [6]

    Devis, N

    N. Devis, N. Demerl \'e , S. Nabi, D. Genova, and P. Esling. Continuous descriptor-based control for deep audio synthesis. In International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023

  7. [7]

    Dhariwal, H

    P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341, 2020

  8. [8]

    Eck and J

    D. Eck and J. Schmidhuber. Finding temporal structure in music: Blues improvisation with lstm recurrent networks. In IEEE workshop on neural networks for signal processing, 2002

Show all 35 references
  1. [9]

    Engel, M

    J. Engel, M. Hoffman, and A. Roberts. Latent constraints: Learning to generate conditionally from unconditional generative models. arXiv preprint arXiv:1711.05772, 2017

  2. [10]

    F. Font, A. P, G. Stegmann, and F. Brinkmann. URL https://github.com/AudioCommons/timbral_models?tab=readme-ov-file. Accessed on 2024-8-16

  3. [11]

    K. K. Ganguli. Guruji padmashree pt. ajoy chakrabarty: As i have seen him. Samakalika Sangeetham, 2012

  4. [12]

    Gopi and F

    S. Gopi and F. William. Introductory studies on raga multi-track music generation of indian classical music using ai. The International Conference on AI and Musical Creativity, 2023

  5. [13]

    Griffin and J

    D. Griffin and J. Lim. Signal estimation from modified short-time fourier transform. IEEE Transactions on acoustics, speech, and signal processing, 32 0 (2): 0 236--243, 1984

  6. [14]

    Gulati, J

    S. Gulati, J. Serr \`a Juli \`a , K. K. Ganguli, S. Sent \"u rk, and X. Serra. Time-delayed melody surfaces for r \=a ga recognition. In Devaney J, Mandel MI, Turnbull D, Tzanetakis G, editors. ISMIR 2016. Proceedings of the 17th International Society for Music Information Ret...

  7. [15]

    Haught-Tromp

    C. Haught-Tromp. The green eggs and ham hypothesis: How constraints facilitate creativity. Psychology of Aesthetics, Creativity, and the Arts, 2017

  8. [16]

    Heitz, L

    E. Heitz, L. Belcour, and T. Chambon. Iterative -(de) blending: A minimalist deterministic diffusion model. In Proc. of the ACM SIGGRAPH Conference, 2023

  9. [17]

    Ho and T

    J. Ho and T. Salimans. Classifier-free diffusion guidance. In NeurIPS Workshop on Deep Generative Models and Downstream Applications, 2021

  10. [18]

    C.-Z. A. Huang, T. Cooijmans, A. Roberts, A. Courville, and D. Eck. Counterpoint by convolution. In International Society for Music Information Retrieval (ISMIR), 2019

  11. [19]

    D. Huron. Sweet anticipation: Music and the psychology of expectation. MIT press, 2008

  12. [20]

    N. A. Jairazbhoy. The r \=a gs of North Indian music: their structure and evolution . Popular Prakashan, 1995

  13. [21]

    Jaques, S

    N. Jaques, S. Gu, R. E. Turner, and D. Eck. Tuning recurrent neural networks with reinforcement learning. In International Conference on Learning Representations, 2017

  14. [22]

    L. Leante. Observing musicians/audience interaction in north indian classical music performance. In Musicians and their Audiences. Routledge, 2016

  15. [23]

    A. McNeil. Seed ideas and creativity in hindustani raga music: beyond the composition--improvisation dialectic. In Ethnomusicology Forum. Taylor & Francis, 2017

  16. [24]

    C. Meng, Y. He, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon. SDE dit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations, 2022

  17. [25]

    Nikrang, K

    A. Nikrang, K. Breckner, T. Neumayr, F. Hirschmann, and M. Augstein. Ai creativity in the light of autonomy. Artificial Intelligence, Co-Creation and Creativity: The New Frontier for Innovation, 2024

  18. [26]

    H. Patel. The Importance of the Teacher , 11 2007. URL https://rhythmicthoughts.wordpress.com/tag/guru-mukhi-vidya/. Accessed on 2024-08-16

  19. [27]

    H. S. Powers and R. Widdess. Theory and practice of classical music. In India, subcontinent of. Grove Music, 2001

  20. [28]

    H. V. Sahasrabuddhe. Analysis and synthesis of hindustani classical music, 1992. URL https://www.cse.iitb.ac.in/ hvs/paper_1992.html. Accessed: 2024-07-22

  21. [29]

    Shikarpur, K

    N. Shikarpur, K. M. Dendukuri, Y. Wu, A. Caillon, and C.-Z. A. Huang. Hierarchical generative modeling of melodic vocal contours in hindustani classical music. In International Society for Music Information Retrieval (ISMIR), 2024

  22. [30]

    Srinivasamurthy, S

    A. Srinivasamurthy, S. Gulati, R. C. Repetto, and X. Serra. Saraga: open datasets for research on indian art music. Empirical Musicology Review, 16 0 (1): 0 85--98, 2021

  23. [31]

    Vidwans, P

    A. Vidwans, P. Verma, and P. Rao. Classifying cultural music using melodic features. In International Conference on Signal Processing and Communications (SPCOM). IEEE, 2012

  24. [32]

    V. Vidwans. Computational music, 2017. URL https://computationalmusic.com/. Accessed on 2024-04-11

  25. [33]

    R. Widdess. Language, Music, and Interaction, chapter Schemas and improvisation in Indian music. College Publications, 2013

  26. [34]

    S.-L. Wu, C. Donahue, S. Watanabe, and N. J. Bryan. Music controlnet: Multiple time-varying controls for music generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024

  27. [35]

    Zhang, A

    L. Zhang, A. Rao, and M. Agrawala. Adding conditional control to text-to-image diffusion models. In IEEE/CVF International Conference on Computer Vision, 2023

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.