{"id":"208d6cbe-d5f6-41bf-a2da-6233c4e4bbf1","arxiv_id":"2608.07376","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper frames unconducted vocal ensemble singing as coupled dynamic systems without an external reference and proposes an architecture for an AI singer that enters, rather than tracks, the ensemble's collective state.","lead":"This paper argues that unconducted vocal ensembles need a new kind of AI music interaction based on continuous, many-to-many coordination instead of call-and-response turns. It lays out a staged research plan and a phone-based rehearsal dataset to test whether an artificial singer can join and reshape an ensemble's collective state, including in rhythms that do not fit a steady beat grid.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"VocalLanes alignment has no independent accuracy baseline; since every downstream collective-state and influence estimate inherits pairwise offsets, the empirical route is not yet established.","rationale":"The reader identified the same weakest assumption, and I agree. The strongest claim is conditional on being able to observe mutual influence at the relevant timescale. Alignment is the only implemented component; representation is described as 'in progress,' and state inference and agent/evaluation are 'proposed.' The paper is honest about the missing independent reference, but that honesty does not make the route empirically testable yet. I do not think this warrants rejection: as a research framework, the problem definition and architecture are coherent, and the alignment gap is a specified engineering validation task rather than a conceptual flaw. The verdict should remain CONDITIONAL, dependent on an independent alignment accuracy evaluation and on specification or implementation of the influence inference. This stress-test does not change the reader's verdict.","tokens_in":6801,"tokens_out":3217,"duration_ms":32288,"concrete_test":"Run a validation session in which 2–3 rehearsals are captured simultaneously with the phone-based VocalLanes setup and a calibrated multichannel recorder (or with a clapperboard/light-sync marker) whose inter-channel delays are known; compare VocalLanes' pairwise offset estimates against ground-truth offsets per singer pair, reporting median absolute error and the fraction of pairs whose estimated lead/lag direction is flipped. If median error exceeds ~20 ms or direction flips are common, the alignment layer cannot support the millisecond-to-second-scale influence inference described in Section 6, and acceptance should hinge on fixing this first.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 reports that pink-noise tests with known offsets met the 40 ms target, and reapplying the detector to two aligned takes found 10–16 ms residual offsets and no drift. The paper then says, 'without an independent reference, this is an internal consistency check rather than an accuracy evaluation.' A minority of real takes approach roughly 30 ms lag that becomes noticeable in rehearsal, although unmeasured. The architecture's scientific claim is that collective-state inference and influence modelling can be tested empirically (Sections 1 and 6), and the only implemented input layer is these pairwise aligned phone recordings. If actual alignment error is 20–30 ms or has sign errors between nearby singers, the proposed temporal-precedence test for influence in A(t)—whether j's recent state improves prediction of i's next event—can produce spurious lead/lag relations. The problem is not that alignment is perfect or imperfect; it is that the error distribution is unknown on real ensemble data, so the load-bearing premise that coordination can be observed at the needed timescale is unverified. This is the weakest point in the chain from field recordings to the claim of an empirically testable route.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that unconducted vocal ensembles cannot be adequately modelled by the call-and-response loop common in interactive music systems, and instead proposes a research framework based on coupled dynamic systems. The central contribution is to define reciprocal coordination without an external reference as a distinct problem and to sketch an empirically testable route to an artificial singer that enters the ensemble's collective state. The manuscript presents Eq. (1) as a minimal formulation of singer state evolution, distinguishes structural, individual, relational, and collective state, treats non-isochronous ritual songs as a hard case, describes the VocalLanes corpus of 40 aligned phone-recorded takes, and proposes a staged pipeline covering representation, state inference, agent policy, and in-situ evaluation. The paper is transparent that influence modelling, agent policy, and comparative evaluation remain proposed, and that the alignment validation performed so far is only an internal consistency check.","tokens_in":6989,"tokens_out":3082,"duration_ms":29563,"significance":"If the proposed research route is pursued successfully, the paper would broaden the design space of human–AI musical interaction beyond turn-taking and reference-based synchronisation, and would provide a computational handle on phenomena such as shared drift, distributed leadership, and non-isochronous temporal contours. The paper's use of a real field corpus, its explicit staging of implementation status, and its attention to consent, cultural specificity, and the limitations of alignment validation are strengths. There are no fitted parameters or quantitative predictions here, so the usual circularity concerns about fitted models do not apply; the paper is best read as a problem definition and a research agenda rather than as an empirical claim. Its main scientific value lies in the falsifiable predictions it sets out for future work, especially the held-out prediction test for mutual influence and the comparative rehearsal study.","major_comments":[{"comment":"The empirical route depends on pairwise phone alignment being accurate enough to resolve coordination at the required timescale, but the only validation reported is a pink-noise test with known offsets and an internal-consistency recheck of two takes with residual offsets of 10–16 ms. The paper itself states that 'without an independent reference, this is an internal consistency check rather than an accuracy evaluation,' and that a minority of real takes approach a roughly 30 ms lag that becomes noticeable in rehearsal. Unknown alignment-error distribution on real ensembles can produce spurious lead/lag relations in the temporal-precedence test for A(t) in Section 6, so the load-bearing premise that coordination is observable at the needed timescale is not yet established. The authors should either add an independent accuracy evaluation (for example, a small set of takes recorded with a simultaneous multi-microphone or click-track reference, with per-pair error distributions reported) or explicitly narrow the empirical testability claim until such validation exists.","section":"Section 4"},{"comment":"The proposed influence criterion — that singer j's recent state improves prediction of singer i's next event beyond q(t), shared drift, and i's history — is not specified enough to be falsifiable as stated. Section 2 correctly notes that signal-derived directionality is not causal ground truth and that the same lag may reflect shared song knowledge or response to a third singer. The paper therefore needs to define the prediction task, the model family, the baselines, the cross-validation protocol, and how common response to unobserved context is excluded. Without this, the A(t) estimates are not uniquely identified, and the planned technical evaluation cannot distinguish reciprocal influence from shared structure.","section":"Section 6"}],"minor_comments":[{"comment":"The delay term d in Eq. (1) is introduced as 'perceptual and system delay' without specifying its units, whether it is a scalar or per-pair, or how it is estimated; please clarify.","section":"Section 2, Eq. (1)"},{"comment":"The caption 'VocalLanes aligns phone recordings for mixed or foreground listening' does not explain the visual encoding of the figure, so a reader cannot tell what the screenshot demonstrates; a sentence describing the interface elements would help.","section":"Figure 1"},{"comment":"The corpus description reports the number of takes and hours but not the distribution of takes over formations or the duration of individual takes; stating the range of take lengths would make the corpus description more complete.","section":"Section 4"},{"comment":"The sentence 'Neither assuming clean stems nor scaling a generic speech or music model resolves the target task' is a strong claim that would benefit from a concrete illustration or a citation to a failed attempt, even though the surrounding discussion of bleed and singing-specific phonetics is reasonable.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its status as a research agenda, which is appropriate for a companion paper. The main risk is that the empirical testability claim in Sections 1 and 6 is currently underwritten by alignment validation that the authors themselves describe as an internal consistency check; this is the point that needs the most work. The self-referential use of VocalLanes and VocalNotes is acceptable given the paper's framing, but an independent validation of the alignment layer would materially strengthen the route to the stated scientific claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look. This short paper does something useful: it names an interaction topology that most musical-agent work ignores—unconducted vocal ensembles where no clock, score, conductor, or leader supplies the reference—and treats non-isochronous ritual song as the hard case rather than an afterthought. The staged architecture (capture, representation, inference, agent/evaluation) is sensible, and the paper is honest that influence, agent policy, and evaluation are proposed rather than done. The discussion of why a beat grid fails for non-isochronous contours is clear, and the citation pattern is relevant without being bloated.\n\nThe soft spots are where the reader puts them, plus one you should keep in view. Eq. (1) is schematic; no inference algorithm is specified; and the empirical route rests on VocalLanes pairwise alignment, whose only validation is pink-noise tests plus an internal recheck. The paper itself says 'without an independent reference, this is an internal consistency check rather than an accuracy evaluation.' That is not a hidden flaw—it is stated—but it means the load-bearing premise, that coordination can be observed at the needed timescale, is unverified on real takes. A minority of outputs approach the roughly 30 ms lag that becomes noticeable in rehearsal, so the error distribution on real ensemble data is genuinely unknown. If pairwise offsets are 20–30 ms or have sign errors, the proposed temporal-precedence influence test is vulnerable to spurious lead/lag relations.\n\nThat said, I would not call this a load-bearing flaw for a framework paper. The authors do not oversell; they explicitly scope the contribution as defining the problem and setting a route. What is missing is not a repair to the argument but a piece of future work: an independent accuracy evaluation of the alignment on real ensemble data, and a specification of the influence inference.\n\nThe self-referential parts did not bother me much. VocalLanes is the corpus under construction and VocalNotes is cited for boundary ambiguity; that is natural. The dataset is not public, which limits reproducibility, but this is a system paper with no fitted parameters, so the circularity burden is low.\n\nWho it is for: people working on musical agents, real-time interactive systems, and ethnomusicology/MIR. It deserves a serious referee; a good referee would ask for the alignment accuracy evaluation and more detail on the inference, not for a completed model. Bring it to reading group if you are thinking about interaction topologies beyond call-and-response.","headline":"A genuinely new problem framing for musical agents—unconducted vocal ensembles without a shared clock—with an honest staged architecture, but the empirical route currently rests on alignment accuracy that has not been independently validated.","tokens_in":7510,"tokens_out":2058,"would_cite":true,"duration_ms":18368,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that unconducted vocal ensembles coordinate through many-to-many reciprocal adjustment with no external clock or tuning reference, making call-and-response the wrong base model, and that an artificial singer can be built…","keywords":["reciprocal coordination","vocal ensembles","human-AI interaction","collective state","non-isochrony","musical agents","audio alignment","coupled dynamic systems"],"falsifier":"An independent reference recording of the same rehearsals would settle it: if a substantial fraction of VocalLanes pairwise alignments deviate by more than roughly 30–40 ms, the collective-state inference step has no reliable input and the proposed route collapses.","tokens_in":6593,"feed_emoji":"🎤","tokens_out":6746,"duration_ms":53515,"temperature":0.7,"pith_summary":"This paper argues that unconducted vocal ensembles are not well modelled as call-and-response loops, because singers affect one another continuously and in many directions with no conductor, score, metronome, or tuning source to fix timing and pitch. It defines reciprocal coordination without an external reference as a distinct problem, and proposes a research architecture in which an artificial singer enters the ensemble's evolving collective state instead of reacting to one human input stream. The architecture runs from field capture of singer-dominant recordings, through singing-aware representation and inference of shared drift and mutual influence, to an agent whose timing policy is conditioned on that inferred state. Non-isochronous ritual songs, whose temporal contour cannot be reduced to a beat grid, are treated as the hard test case for the framework.","feed_headline":"Call-and-response is the wrong base model for vocal ensembles","feed_subtitle":"An AI singer can enter an ensemble's shared timing and mutual pull instead of merely replying to a lead.","key_machinery":"The load-bearing machinery is the coupled-state equation and the influence network $A(t)$ it introduces, together with the distinction between structural, individual, relational, and collective state. Structural state locates the performance within a learned song; individual state tracks each singer's timing, pitch, and uncertainty; relational state captures pairwise lag and correction; collective state describes higher-order organisation such as common drift or distributed leadership. For non-isochronous repertoire the paper replaces the beat grid with a probabilistic verse-shape built from hierarchical semi-Markov or Markov-renewal models, so periodicity is learned as a property of repertoire rather than imposed as a prerequisite. VocalLanes supplies the multi-channel, singer-dominant, repeated-take recordings on which inference of these states would run.","core_discovery":"The central claim is that collective organisation in unconducted singing emerges from many-to-many reciprocal adjustment, so the proper computational object is not a tempo estimate or a synchrony score but the changing structure of mutual influence among participants. The paper formalizes each singer's next state as $x_i(t+1)=f_i(x_i(t), x_{-i}(t-d), A_i(t), q(t), c)$, where $A_i(t)$ holds time-varying influence weights from other singers, $q(t)$ is a learned structural representation of the current song and phrase, and $c$ bundles singer-, genre-, and tradition-specific constraints; no term is a privileged global clock. On this view a performance can remain highly coordinated while absolute tempo and pitch drift, because coherence is relational. The paper's stated contribution is to define reciprocal coordination without an external reference as a distinct problem and set out an empirically testable route to an artificial singer within the ensemble's dynamics, with the VocalLanes corpus of aligned phone recordings as the data foundation.","pith_inferences":["The same relational-coordination formalism could be carried into spoken conversation, dance, sports, or any joint action where timing is co-produced without a privileged reference; the paper does not develop these applications.","A direct perturbation test would be to have the agent deliberately pull or hold a phrase while human singers continue; if singers measurably shift their timing in response, the influence network $A(t)$ is being manipulated and can be checked against the singers' own accounts.","The hardest unresolved dependency is data quality: unless pairwise phone alignment in real rehearsals is shown against an independent reference to stay near the 40 ms target, the proposed collective-state inference has no verified input no matter how sound the model is."],"forward_implications":["If the formulation is right, interactive music systems for ensemble singing should be judged by how they change mutual influence among singers, not by how accurately they follow a beat or a score.","A coordinated artificial singer can be built without assuming a global clock: its influence can be graded by uncertainty, so it may join a shared phrase expansion, reduce its pull while leadership is unresolved, or temporarily follow.","The framework implies that high coordination and large departure from an external tempo or tuning standard can coexist, making absolute alignment metrics the wrong success criterion.","Non-isochronous ritual songs become the hard test case: they require a model that knows where the ensemble is within a learned structural contour before 'ahead' or 'behind' has any meaning."],"supporting_citations":[{"why":"supplies the precedent of interpersonal synergies, higher-order systems in which participants are coupled and compensate for one another rather than execute independent plans.","marker":"[21]"},{"why":"supplies the dynamics-of-attending account used to justify probabilistic verse-shapes rather than a fixed beat grid.","marker":"[12]"},{"why":"supplies the ethnomusicological concept of entrainment that the relational-coordination view extends beyond periodic time.","marker":"[4]"},{"why":"supplies the DTW alignment method reserved for unreliable pairwise offsets in VocalLanes.","marker":"[8]"},{"why":"shows that even expert note-boundary judgments vary across musical cultures, motivating culturally situated alignment rather than assumed onsets.","marker":"[20]"},{"why":"shows phoneme-boundary alignment is less precise in singing than speech, motivating the singing-aware representation stage.","marker":"[26]"},{"why":"supplies the bespoke HSMM extension for real-time singing-voice score following that the proposed agent pipeline builds on.","marker":"[10]"},{"why":"documents the small size of existing choir multitrack datasets, motivating the new VocalLanes corpus.","marker":"[7]"}],"fun_headline_variants":["AI singer that joins the ensemble's mutual pull, not just replies","Reciprocal coordination, not call-and-response, in AI vocal ensembles","AI singer enters the ensemble's shared timing, not merely answers","From call-and-response to shared timing for AI singers","AI singer that adjusts, not answers, in vocal ensembles"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire pipeline depends on the phone-recorded pairwise alignments being accurate enough — the paper's own target is 40 ms, supported only by pink-noise tests and an internal consistency recheck — because every later stage consumes those alignments as if they were ground truth for coordination.","fun_headline_variants_meta":{"raw":{"variants":["AI singer that joins the ensemble's mutual pull, not just replies","Reciprocal coordination, not call-and-response, in AI vocal ensembles","AI singer enters the ensemble's shared timing, not merely answers","From call-and-response to shared timing for AI singers","AI singer that adjusts, not answers, in vocal ensembles"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001059,"raw_usage":{"total_tokens":4441,"prompt_tokens":942,"completion_tokens":3499,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":3413}},"tokens_in":558,"tokens_out":3499,"duration_ms":19224,"temperature":1.0,"reasoning_tokens":3413,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T05:23:40.154752+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent reference recording of the same rehearsals would settle it: if a substantial fraction of VocalLanes pairwise alignments deviate by more than roughly 30–40 ms, the collective-state inference step has no reliable input and the proposed route collapses.","supporting_citations":[{"cited_title":"Large and Mari Riess Jones","cited_arxiv_id":null,"evidence_quote":"supplies the dynamics-of-attending account used to justify probabilistic verse-shapes rather than a fixed beat grid."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the ethnomusicological concept of entrainment that the relational-coordination view extends beyond periodic time."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the DTW alignment method reserved for unreliable pairwise offsets in VocalLanes."},{"cited_title":"2025–2026","cited_arxiv_id":null,"evidence_quote":"shows that even expert note-boundary judgments vary across musical cultures, motivating culturally situated alignment rather than assumed onsets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"shows phoneme-boundary alignment is less precise in singing than speech, motivating the singing-aware representation stage."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the bespoke HSMM extension for real-time singing-voice score following that the proposed agent pipeline builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"documents the small size of existing choir multitrack datasets, motivating the new VocalLanes corpus."}],"review_version":1}