{"id":"19195227-1fd5-467f-9444-265e7c62e211","arxiv_id":"2608.05257","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Composers of eighteenth-century opera seria encoded character gender in vocal writing but barely tailored notation to the singer's biological sex.","lead":"This paper uses statistical models on 1,682 notated arias to test whether composers wrote differently for masculine versus feminine characters and for castrato versus female sopranos. It finds that character gender is measurably encoded in the vocal writing, while the singer's sex leaves only a weak trace.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim is not identified because singer sex and character gender are 86% coincident and no model includes both covariates, so the 'dramatic rather than physiological' conclusion overreaches.","rationale":"Label noise in the name-based sex inference is a valid secondary concern, and the reader correctly identified it as a stated limitation. However, a sensitivity analysis of label noise would not, by itself, establish the paper's central causal claim, because the 86% correlation between the two target variables means that separate marginal classifiers cannot separate 'writing for the character' from 'writing for the singer.' Even with perfectly accurate sex labels, the observed asymmetry could arise from the label correlation; the cross-cast analysis directly tests the partial associations. Since the paper's headline contribution is the dramatic-not-physiological interpretation, the confounding and the missed cross-cast analysis is the single most load-bearing concern. The verdict remains CONDITIONAL because the required analysis is feasible with the deposited data and code.","tokens_in":22224,"tokens_out":12167,"duration_ms":111572,"concrete_test":"Fit the Case study 1 character-gender classifier (or retrain a ridge logistic regression on the full corpus) and evaluate it on the 197 cross-cast arias where singer sex and character gender diverge. If accuracy on this subset is at or above the held-out accuracy in Table 3 (about 0.61), character gender is encoded independently of singer sex. Additionally, fit a ridge logistic regression predicting singer sex that includes character gender as a covariate; if the singer-sex coefficients become non-significant while character-gender coefficients persist, the asymmetry is supported. Report the cross-cast sample sizes and bootstrap confidence intervals.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is not label noise but the confounded comparison underlying the causal conclusion. Singer sex and character gender are coincident in 86.28% of the 1,436 arias with known singers (Section 3.4). Case study 1 predicts character gender from musical features; Case study 3 predicts singer sex from the same features. The paper interprets the lower accuracy in Case study 3 as evidence that composers encoded the character rather than the singer. However, no model includes both covariates. The claimed gradient across Tables 7–8—e.g., HighestNoteIndex odds decrease of 20.43% (CS1), 15.42% (CS2), 7.78% (CS3)—is estimated on different subsets and does not identify the partial association of each label with the music. The abstract's assertion that 'female singers sang higher pitches than male sopranos, but this correlates more with character portrayal than with physiology' is a causal claim about confounding, but no regression with both character gender and singer sex as predictors is reported. The conclusion that gender coding is 'dramatic rather than physiological in origin' therefore overreaches the evidence. The corpus contains 197 cross-cast arias (74 masculine characters sung by female sopranos; 123 feminine characters sung by male sopranos, per Table 2) that could directly test the claim; the authors do not analyze them.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper analyzes a corpus of 1,682 soprano arias from settings of Metastasio’s five most popular drammi per musica (1724–1810). Using musical features extracted from the soprano part, the authors compare ridge and lasso logistic regression and random forest classifiers for three tasks: predicting the gender of the dramatic character (Case study 1), predicting character gender within a subsample of arias where singer sex and character gender align (Case study 2), and predicting the singer’s inferred biological sex (Case study 3). The paper’s central claim is that character gender is more strongly encoded in the notated vocal writing than singer sex, so that gender coding in opera seria was ‘dramatic rather than physiological in origin.’ The reported held-out accuracies are modest (0.58–0.64) but significantly above the majority-class baseline for the main models, and coefficient interpretations identify higher maximum pitches, smaller intervals, and lower intervallic variability with feminine/female classes, and larger ambitus, larger leaps, and higher vocal presence with masculine/male classes.","tokens_in":22444,"tokens_out":8973,"duration_ms":73331,"significance":"If the central claim could be rigorously established, the study would provide important quantitative evidence on a long-debated musicological question, showing that composers encoded gender in the musical text rather than relying on vocal type, and that the documented interchangeability of castrati and female sopranos coexisted with strongly gendered vocal writing. The paper deserves credit for a large, carefully curated corpus; explicit held-out evaluation with bootstrap significance tests; Bonferroni-corrected asymptotic confidence intervals for ridge coefficients; and public deposition of data and code. The main results are transparent with respect to apparent vs. held-out metrics. The significance of the work, however, depends on whether the authors can support the causal interpretation that character gender drives the musical differences beyond the confounded singer-sex signal.","major_comments":[{"comment":"The conclusion that composers encoded gender ‘dramatically rather than physiologically’ (Section 5) is not identified by the reported models. Singer sex and character gender coincide in 86.28% of the 1,436 arias with known singers (Section 3.4, Table 2), and the manuscript itself states that ‘this case study cannot fully disentangle the one from the other.’ Yet no model includes both covariates: Case study 1 predicts character gender, Case study 3 predicts singer sex, and the comparison of their accuracies is across different subsets with different sample sizes. The lower held-out accuracy in Case study 3 (0.585 vs. 0.609 in Case study 1) is entirely compatible with a null model in which only character gender matters and singer sex has no independent predictive effect, given the 86% alignment. To support the paper’s causal claim, the authors should either (a) fit a joint model for singer sex that includes character gender as a covariate (or the converse), or (b) analyze the 197 cross-cast arias (74 masculine roles sung by female singers and 123 feminine roles by male singers, Table 2) and test whether, within a fixed character gender, musical features differ by singer sex. Without such an analysis, the abstract’s statement that ‘Female singers sang higher pitches than male sopranos, but this correlates more with character portrayal than with physiology’ overreaches the evidence.","section":"Section 3.4, Section 5"},{"comment":"The interpretation of the feature gradients as gender-coded writing is also confounded with dramaturgical variables such as role rank and affective content. The authors themselves note in Section 4 that ‘pity’ arias are 165 for feminine characters vs. 113 for masculine ones while ‘anger’ arias are 194 for masculine vs. 31 for feminine, and that the association of VoicePresence and AverageDuration with masculine characters may reflect ‘differences in dramatic prominence’ rather than gender. Since the models include no control for role rank, aria type, or emotion, the conclusion that the identified features encode character gender rather than these correlated dramaturgical dimensions is not yet established. At minimum, the causal framing in Sections 4 and 5 should be tempered, or an analysis that adjusts for these variables (or reports the sensitivity of the coefficients to their inclusion) should be added.","section":"Section 4"},{"comment":"The singer-sex labels are inferred solely from given names in libretti, a procedure the authors themselves flag as ‘susceptible to inaccuracies’ (Section 2.1). Since the central negative result—that singer sex is only weakly recoverable from the music—depends entirely on the validity of these labels, a sensitivity analysis is needed. The authors should, for example, exclude arias whose singers’ names are ambiguous, cross-check a subsample against documented castrato biographies or the CORAGO database, and rerun the Case study 3 classification to demonstrate that the weak signal is not an artifact of label noise. As it stands, the one-sentence acknowledgment of the limitation is not commensurate with the load-bearing role that the singer-sex labels play in the paper’s central claim.","section":"Section 2.1, Section 3.4"}],"minor_comments":[{"comment":"The notation β−1 is used without definition; it should be explicitly introduced as the coefficient vector β with the intercept β0 removed.","section":"Section 2.4, Eq. (1)"},{"comment":"The parenthetical explanation of LargestSemitonesDesc is difficult to parse; it should be rewritten to state clearly that descending intervals are encoded as negative values, so a larger numeric value means a smaller absolute leap.","section":"Section 3.4"},{"comment":"The features are not listed in the order in which they are discussed in Sections 3.2–3.3; reordering the rows (e.g., by the Case study 1 effect size) would improve readability.","section":"Table 7"},{"comment":"The statement that the bootstrap ‘does not retrain the models M1 and M2’ is useful, but the text should explicitly note that the resulting intervals therefore cover testing-set variability only, not the variability induced by hyperparameter selection or the CV-based selection of λ̂1SE.","section":"Section 2.6"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The manuscript is statistically careful in its implementation, but the central inference is not yet supported by the design: the comparison between Case studies 1 and 3 conflates the two covariates. The authors do acknowledge the 86.28% overlap in Section 3.4, but they do not act on it. The required additional analysis (a joint model or a cross-cast sub-analysis) is straightforward and within the scope of a revision. The paper may also be on the musicology-heavy side for a statistics journal; however, if the journal accepts applied data analysis with substantive interpretation, the topic is suitable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading: this is the first large-scale corpus study of soprano vocal typology in opera seria, built on 1,682 arias with deposited data and code. The methodological care is real: stratified held-out splits, bootstrap significance tests, Bonferroni-corrected confidence intervals, model comparison, and a clear distinction between apparent and held-out metrics. The authors also flag the main label-noise concern themselves — singer sex is inferred from given names. Credit where due: this is a solid, reproducible piece of empirical musicology.\n\nThe soft spots are real but addressable. The biggest one is the confounding issue. Singer sex and character gender coincide in 86.28% of the 1,436 arias with known singers, yet no model includes both covariates. Case study 1 predicts character gender; case study 3 predicts singer sex; the lower accuracy in case study 3 is then interpreted as evidence that composers encoded the character rather than the singer. That inference is not identified. The 197 cross-cast arias — 74 masculine characters sung by women, 123 feminine characters sung by male sopranos — are the natural place to test the claim, and they are not analyzed. The abstract's statement that higher pitches in female singers 'correlates more with character portrayal than with physiology' overreaches what a single-label predictive model can show. The conclusion that gender coding was 'dramatic rather than physiological in origin' is a causal claim about a confounded comparison.\n\nA second, related issue: the label noise in case study 3 is acknowledged but not quantified. A sensitivity analysis — e.g., dropping ambiguous names, or re-running with only unambiguous cases — would strengthen the null result. The final Table 6 apparent metrics are clearly labeled as such, so that is minor.\n\nWho benefits: computational musicologists, opera scholars, and anyone working on gender and musical corpora. It deserves a serious referee, but the referee should push for a joint model (both character gender and singer sex as predictors) or at least a focused analysis of the cross-cast subset, and a tempered conclusion. I would engage with it, and I would cite the corpus and the descriptive results, but I would not cite the causal framing.\n\nRecommendation: send to peer review, with the expectation of substantial revision on the confounding issue.","headline":"A careful, reusable corpus study whose central interpretive claim about dramatic vs. physiological gender coding outruns the statistics, because singer sex and character gender are never jointly modeled.","tokens_in":22999,"tokens_out":1290,"would_cite":true,"duration_ms":15045,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that in eighteenth-century Italian opera seria, composers encoded the dramatic gender of characters in the melodic fabric of soprano arias, while the biological sex of the singer is only weakly recoverable from the…","keywords":["opera seria","castrato","soprano","vocal typology","gender in music","corpus study","ridge logistic regression","statistical learning"],"falsifier":"A decisive check would be to obtain archival documentation of the actual sex of the 303 premiere singers, from payrolls, chapel records, or castrato contracts, instead of inferring it from given names, and rerun Case study 3 on the same features. If held-out accuracy then rises well above the 0.585 ridge result toward the 0.61-0.64 range seen for character gender, or if a physiology-linked feature such as long-note duration or coloratura density becomes significantly predictive while character gender is held constant, the paper's asymmetry claim would be in doubt.","tokens_in":1798,"feed_emoji":"🎭","tokens_out":2157,"duration_ms":69682,"temperature":0.7,"pith_summary":"Eighteenth-century opera seria was written overwhelmingly for sopranos, and the same high register was filled by both castrati and women, often performing roles of the opposite gender. The paper asks whether composers tailored their vocal writing to the singer's sex or to the character's gender, using about 1,700 surviving arias from the five most popular Metastasio libretti. The central finding is an asymmetry: character gender is consistently recoverable from the music, though only modestly above chance, while singer sex is barely recoverable. Masculine characters get wider ranges, larger and more variable intervals, and more vocal presence; feminine characters get higher peaks, smaller and minor intervals, and steadier lines. The authors conclude that gender coding in opera seria was dramatic rather than physiological, anchored to the character rather than the body that sang it.","feed_headline":"Composers wrote character gender into arias; singer sex barely shows","feed_subtitle":"A corpus of 1,682 arias shows masculine and feminine roles diverge musically while male and female singers' parts barely do.","key_machinery":"The machinery is a corpus of 1,682 symbolic scores of soprano arias from the five most frequently reset Metastasio drammi, reduced to 86 vocal-part features covering range, highest and lowest notes, interval sizes and qualities, note durations, and vocal presence, then preprocessed to 34 or 35 standardized predictors. The chosen model is ridge logistic regression, a regularized classifier whose coefficients estimate how each standardized feature shifts the odds of the masculine/male class and come with asymptotically valid, Bonferroni-corrected confidence intervals; it is selected because its accuracy is statistically comparable to lasso and random forest while allowing interpretation and uncertainty quantification. The classification targets are character gender in two case studies and the singer's sex inferred from given names in libretti in a third. The decisive quantities are the standardized coefficients, and comparing these coefficients across case studies is what reveals the asymmetry.","core_discovery":"The paper's core discovery is that gender was written into the notes, while the singer's body was not. Across three binary classification case studies on 1,682 soprano arias, ridge logistic regression identifies significant, interpretable associations between musical features and the character's gender, with held-out accuracy around 0.61 for character-gender classification and around 0.60 when only cast-aligned arias are used, both above the majority-class baseline. The same model applied to the premiering singer's inferred biological sex attains only about 0.58, with weaker coefficients and only a modest edge over the baseline. Specific features tell the same story: the highest note of the aria loses more than half of its predictive weight when the target changes from character gender to singer sex, and the interval-size features that flag masculinity also weaken. The paper reads this as evidence that the flexibility long documented in casting practice was sustained by sharply gendered composition, and that the celebrated castrato difference was a difference in sound, not in notated notes.","pith_inferences":["One consequence the authors do not spell out: if gender coding is dramatic rather than physiological, historically informed staging need not treat these arias as body-locked; casting decisions could follow the dramatic characterization rather than a voice-type requirement.","A direct perceptual test of the written-code claim would be to play paired excerpts to listeners and ask them to guess character gender versus singer sex; the theory predicts a large gap in accuracy, mirroring the classification gap.","Because 86% of the known-singer arias align character gender with singer sex, the weak case-study-3 signal may partly reflect label noise from name-based sex inference; a cross-cast-only analysis, though small, would be the sharper comparison and could be run with the deposited dataset.","The paper's framing suggests a general technique for historical performance conventions with interchangeable bodies: comparing classification accuracy between role-level and performer-level targets isolates which variable the notation actually encodes."],"forward_implications":["If the conclusion is right, claims that opera seria register carried no dramatic connotation must be revised: gender marking demonstrably sits in the melodic fabric.","The documented flexibility of casting did not imply neutral writing; composers wrote gender-coded lines while planning for a market in which either sex might take a role.","The weak singer-sex signal indicates that the castrato versus female-soprano distinction, so vivid to contemporaries, was primarily acoustic and performative, and is largely invisible in score-based features.","Features like vocal presence and average note duration attach to the character rather than the singer, suggesting that endurance and long-held notes belonged to heroic masculinity as a dramatic trait, not to castrato physiology.","The asymmetry motivates future work with explicit tessitura descriptors and with modeling of dramatic rank, which the authors identify as beyond the present scope."],"supporting_citations":[{"why":"Supplies the historical documentation of castrato range, coloratura, long-held notes, and cross-casting that motivates the voice-specific writing hypothesis.","marker":"Seedorf, 2015"},{"why":"Documents the acoustic uniqueness of the castrato voice and the idea that the notated score does not capture it, which frames the weak singer-sex signal.","marker":"Feldman, 2015"},{"why":"Provides the DIDONE arias database with the scores, characters, and premiere singer metadata that constitute the corpus.","marker":"Llorens et al. (2024)"},{"why":"Supplies the musif feature-extraction software used to derive the 86 musical features from the MusicXML scores.","marker":"Llorens et al. (2023)"},{"why":"Provides the lasso, ridge, and elastic-net statistical learning framework within which the models are fitted.","marker":"Hastie et al. (2015)"},{"why":"Supplies the asymptotic covariance formula for ridge logistic regression used to build the confidence intervals on coefficients.","marker":"Le Cessie and Van Houwelingen (1992)"},{"why":"Provides the libretto metadata, including the singer names from which biological sex is inferred.","marker":"Pompilio (2025)"},{"why":"Documents the role hierarchy and tenor/authority axis of opera seria, grounding the interpretation of vocal presence and range as dramatic markers.","marker":"Torrente and Domínguez (2025)"}],"fun_headline_variants":["Soprano arias reveal character gender, not singer sex","In 1700s opera, the role's gender is in the notes; the singer's is not","Character gender is composed; singer sex is barely there","In opera seria, the music knows the character's gender, not the singer's"],"cache_read_input_tokens":25088,"weakest_assumption_plain":"The paper assumes that a singer's biological sex can be inferred from whether their given name in a libretto is traditionally male or female, and the authors themselves flag that this method is susceptible to inaccuracies; if many labels are wrong, the case-study-3 conclusion that singer sex is barely recoverable could be an artifact of noisy labels rather than a property of the music.","fun_headline_variants_meta":{"raw":{"variants":["Soprano arias reveal character gender, not singer sex","In 1700s opera, the role's gender is in the notes; the singer's is not","Character gender is composed; singer sex is barely there","In opera seria, the music knows the character's gender, not the singer's"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000504,"raw_usage":{"total_tokens":2465,"prompt_tokens":954,"completion_tokens":1511,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":1429}},"tokens_in":570,"tokens_out":1511,"duration_ms":11112,"temperature":1.0,"reasoning_tokens":1429,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T16:55:51.846548+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check would be to obtain archival documentation of the actual sex of the 303 premiere singers, from payrolls, chapel records, or castrato contracts, instead of inferring it from given names, and rerun Case study 3 on the same features. If held-out accuracy then rises well above the 0.585 ridge result toward the 0.61-0.64 range seen for character gender, or if a physiology-linked feature such as long-note duration or coloratura density becomes significantly predictive while character gender is held constant, the paper's asymmetry claim would be in doubt.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the historical documentation of castrato range, coloratura, long-held notes, and cross-casting that motivates the voice-specific writing hypothesis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the acoustic uniqueness of the castrato voice and the idea that the notated score does not capture it, which frames the weak singer-sex signal."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the musif feature-extraction software used to derive the 86 musical features from the MusicXML scores."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the lasso, ridge, and elastic-net statistical learning framework within which the models are fitted."},{"cited_title":"and Van Houwelingen , H","cited_arxiv_id":null,"evidence_quote":"Supplies the asymptotic covariance formula for ridge logistic regression used to build the confidence intervals on coefficients."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the libretto metadata, including the singer names from which biological sex is inferred."},{"cited_title":"and Dom \\'i nguez , J","cited_arxiv_id":null,"evidence_quote":"Documents the role hierarchy and tenor/authority axis of opera seria, grounding the interpretation of vocal presence and range as dramatic markers."}],"review_version":1}