Pith. sign in

REVIEW 2 major objections 2 minor 10 references

Two Roman numeral analysis collections are reconciled into one interoperable corpus that keeps 84 pieces under both original annotations for direct comparison.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-07-01 02:52 UTC pith:MT7ECBWT

load-bearing objection The paper merges two Roman numeral datasets into 1621 pieces with 84 dual-annotated overlaps, but the transformations lack reported quantitative checks on fidelity. the 2 major comments →

arxiv 2606.31595 v1 pith:MT7ECBWT submitted 2026-06-30 cs.SD cs.DLeess.AS

Dilemmadata: On the Interoperability of Heterogeneous Roman Numeral Datasets

classification cs.SD cs.DLeess.AS
keywords Roman numeral analysismusic datasetsdata interoperabilitytonal harmonycorpus creationannotation reconciliationmusic information retrieval
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper creates dilemmadata by converting two separate collections of Roman numeral chord labels into a single shared TSV format. It confronts differences in annotation vocabulary, data representation, software tools, and piece selection by transforming or omitting information while recording those choices. The result is a merged set of 1,621 pieces containing roughly 2.8 million note-wise annotations, including an overlap of 84 pieces that can be examined under each original analytical standard. A sympathetic reader would care because the shared reference set makes it possible to test how two legitimate ways of labeling the same music affect downstream modeling tasks in tonal harmony.

Core claim

By deliberately transforming, augmenting, and omitting information to resolve annotation-standard, representational, toolchain, and curatorial mismatches, the two source collections become interoperable under a common note-wise TSV schema, yielding dilemmadata as the largest homogeneous Roman-numeral corpus currently available while retaining 84 common pieces under each of their original analyses.

What carries the argument

The shared note-wise TSV schema that formalizes mismatches in vocabulary, syntax, chord extensions, and special functions while flagging transformations that may affect fidelity.

Load-bearing premise

The transformations and omissions needed to reconcile the two datasets preserve enough musical semantics that the merged annotations remain valid for comparative modeling.

What would settle it

An independent expert audit of the 84 overlapping pieces that checks whether each converted annotation still matches the musical intent of its source corpus.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Comparative modeling can now train and evaluate systems on identical musical material labeled under two different analytical traditions.
  • Consistency checks across the overlap provide a preliminary basis for critiquing the theoretical assumptions in each original standard.
  • The unified corpus supports larger-scale data-driven studies in tonal harmony than either source alone.
  • Future work can refine Roman-numeral encoding standards by building on the documented reconciliation rules.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The overlap set could serve as a benchmark for measuring how much annotation differences affect automatic harmonization accuracy.
  • Extending the same reconciliation approach to additional corpora might produce even larger multi-tradition resources.
  • The flagging of subtle fidelity changes offers a template for auditing other music-annotation merges.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The manuscript introduces dilemmadata as the first reconciled resource merging the AugmentedNet Dataset (AN) and Distant Listening Corpus (DLC) into a shared note-wise TSV schema. It describes deliberate transformations to resolve four families of dilemmas (annotation-standard differences in vocabulary/syntax/extensions/special functions, representational choices, toolchain incompatibilities between music21 and ms3+dimcat, and curatorial decisions on inclusion/exclusion), while claiming to formalize mismatches and preserve musical semantics. The resulting corpus contains 1,621 pieces and approximately 2.8 million annotations; crucially, 84 shared pieces are retained with both original analyses to enable note-for-note comparison of analytical traditions.

Significance. If the semantic-preservation claim holds, the work removes a practical interoperability barrier in computational musicology and supplies the largest homogeneous Roman-numeral corpus to date together with a dual-annotated reference set. This directly supports comparative harmonization modeling and future refinement of encoding standards; the Zenodo release further aids reproducibility.

major comments (2)
  1. [Abstract and consistency checks] Abstract and consistency-checks description: the central claim that the merged corpus is 'homogeneous' and that transformations 'preserve musical semantics' rests on unshown evidence; no quantitative metrics (e.g., fraction of annotations requiring non-invertible changes, number of fidelity-loss flags, or agreement rates on the 84 shared pieces) are reported to substantiate the 'preliminary assessment of post-conversion validity'.
  2. [Transformation rules] Transformation-rules section: explicit, reproducible mapping rules for reconciling vocabulary size, syntax, chord extensions, and special functions are not supplied, preventing independent verification that the deliberate augmentations, omissions, and conversions maintain sufficient musical semantics for downstream comparative use.
minor comments (2)
  1. [Abstract] The abstract contains the typo 'aprox.'; correct to 'approx.'.
  2. A summary table listing piece and annotation counts before/after merging, plus counts of transformations per dilemma family, would improve clarity and allow readers to gauge scale of changes.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their constructive comments, which highlight opportunities to strengthen the evidentiary basis and reproducibility of our reconciliation process. We address each major comment below and will incorporate revisions to improve the manuscript.

read point-by-point responses
  1. Referee: [Abstract and consistency checks] Abstract and consistency-checks description: the central claim that the merged corpus is 'homogeneous' and that transformations 'preserve musical semantics' rests on unshown evidence; no quantitative metrics (e.g., fraction of annotations requiring non-invertible changes, number of fidelity-loss flags, or agreement rates on the 84 shared pieces) are reported to substantiate the 'preliminary assessment of post-conversion validity'.

    Authors: We agree that the current version relies on a qualitative description of consistency checks without accompanying quantitative metrics. This is a valid observation. In the revised manuscript we will add a dedicated subsection (likely in Section 4 or 5) that reports: (i) the total number and fraction of annotations undergoing each class of transformation, (ii) counts of fidelity-loss flags raised during conversion, and (iii) note-level agreement statistics computed on the 84 dual-annotated reference pieces. These figures will be derived from the existing consistency-check scripts already used to produce the corpus. revision: yes

  2. Referee: [Transformation rules] Transformation-rules section: explicit, reproducible mapping rules for reconciling vocabulary size, syntax, chord extensions, and special functions are not supplied, preventing independent verification that the deliberate augmentations, omissions, and conversions maintain sufficient musical semantics for downstream comparative use.

    Authors: We acknowledge that the manuscript describes the four families of dilemmas at a high level but does not enumerate the concrete, machine-readable mapping rules in a form that permits direct replication. This limits independent verification. In the revision we will add an explicit mapping table (or supplementary file) that lists, for each original vocabulary element, the target representation, any augmentation or omission applied, and the rationale for semantic preservation. The table will be accompanied by the Python conversion scripts (already used to generate the released Zenodo data) so that the rules become fully reproducible. revision: yes

Circularity Check

0 steps flagged

No circularity: data-curation effort with no derivations or fitted predictions

full rationale

The paper is a schema-mapping and corpus-reconciliation project that describes deliberate transformations to merge two annotation collections into a shared TSV format. It contains no equations, no parameter fitting, no predictions of downstream quantities, and no load-bearing self-citations or uniqueness theorems. The retained 84-piece overlap set and the claim of semantic preservation are direct consequences of the explicit reconciliation steps rather than reductions by construction; the work is therefore self-contained against external benchmarks with no circular steps.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

The central claim rests on the domain assumption that a single note-wise TSV schema plus selective transformations can preserve the musical semantics of two heterogeneous annotation standards sufficiently for comparative use.

axioms (1)
  • domain assumption A shared note-wise TSV schema can represent the essential information from both original standards without critical loss of musical meaning.
    Invoked when the abstract describes resolving annotation-standard and representational dilemmata by transforming and omitting information while preserving semantics.

pith-pipeline@v0.9.1-grok · 5863 in / 1411 out tokens · 73437 ms · 2026-07-01T02:52:50.225252+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Dilemmadata: On the Interoperability of Heterogeneous Roman Numeral Datasets." pith.science (2026). https://pith.science/paper/MT7ECBWT

@misc{pith2026260631595,
  author       = {Pith},
  title        = {Pith review of: Dilemmadata: On the Interoperability of Heterogeneous Roman Numeral Datasets},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MT7ECBWT}},
  note         = {Machine review of arXiv:2606.31595}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In recent years, there has been growing effort to annotate and collect large-scale corpora of Roman numeral analyses in support of data-driven studies in tonal harmony. We introduce dilemmadata, the first resource to reconcile two major collections, the AugmentedNet Dataset (AN) and the Distant Listening Corpus (DLC), making them interoperable through a shared note-wise TSV schema. The reconciliation confronts four families of dilemmata: annotation-standard (the two encode the same musical fact differently in terms of vocabulary size, syntax, conventions for chord extensions, inventory of special chord functions), representational (what counts as a row, and which information survives the conversion), toolchain (incompatible Python ecosystems built around music21 vs. ms3+dimcat), and curatorial (which pieces to include, exclude, or retain twice). We resolve each by deliberately transforming, augmenting, and omitting information, formalising the mismatches, preserving musical semantics, and flagging transformations that may subtly affect annotation fidelity. Consistency checks and qualitative inspections offer a preliminary assessment of post-conversion validity and a basis for critiquing the theoretical assumptions embedded in each original standard. After removing duplicates and merging the two collections, the resulting dilemmadata (1,621 pieces and aprox. 2.8 M note-wise annotations) is the largest homogeneous Roman-numeral corpus currently available, albeit far from perfect. Crucially, we retain 84 pieces common to both corpora under each of their original analyses, yielding a shared reference set in which two equally legitimate analytical traditions can be compared note-for-note over identical musical material. Released on Zenodo, dilemmadata supports interoperability, comparative harmonization modeling, and future refinement of Roman-numeral encoding standards.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

10 extracted references · 10 canonical work pages

  1. [1]

    and the Distant Listening Corpus (DLC) (Hentschel et al., 2025). Although both encode human expert analyses in the form of Roman numeral labels, in practice however, harmonising them into a single training and research corpus proved cumbersome and error-prone due to disparate encoding paradigms. We present the outcome of this effort: dilemmadata , the lar...

  2. [2]

    Data processing pipeline from left to right. The left side depicts the AugmentedNet Dataset and the Distant Listening Corpus as meta-corpora comprising multiple sub-corpora in their own right, one of them identical: the Annotated Beethoven Corpus, ABC, accounting for 70 of the 99 overlapping pieces (remaining overlaps stem from distinct sources and provid...

  3. [3]

    1.2 The Distant Listening Corpus (DLC) The DLC 2 provides 1,283 pieces from four centuries, symbolically encoded in MuseScore’s MSCX format and enriched with embedded DCML harmony annotations, including Roman numerals, key, phrase, and cadence labels, directly integrated as textual objects within the score files (Hentschel et al., 2025). These annotated sc...

  4. [4]

    ground truth

    The three annotation layers show labels and harmonic rhythm from the two source datasets and, in the bottom system, of our respective simplified representation without inversions and secondary keys (which are represented elsewhere in our dataset). 2 Dataset Alignment 2.1 Disparate encoding paradigms The AND relies on the RomanText encoding standard and par...

  5. [5]

    https://doi.org/10.5281/zenodo.10265338 Hentschel, J., Moss, F

    , 516–523. https://doi.org/10.5281/zenodo.10265338 Hentschel, J., Moss, F. C., McLeod, A., Neuwirth, M., & Rohrmeier, M. (2022). Towards a Unified Model of Chords in Western Harmony. In S. Münnich & D. Rizo (Eds.), Music Encoding Conference Proceedings 2021 (pp. 143–149). https://doi.org/10.17613/5rxtc-wcj65 Hentschel, J., Moss, F. C., Neuwirth, M., & Rohr...

  6. [6]

    https://doi.org/10.1038/s41597-025-04976-z Hentschel, J., & Rohrmeier, M. (2023). ms3: A parser for MuseScore files, serving as data factory for annotated music corpora. Journal of Open Source Software , 8 (88),

  7. [7]

    https://doi.org/10.21105/joss.05195 Huron, D. (2020). **harm Representation for Western Functional Harmony . Humdrum Representations. https://www.humdrum.org/rep/harm Karystinaios, E., & Widmer, G. (2023). Roman Numeral Analysis With Graph Neural Networks: Onset-Wise Predictions From Note-Wise Features. Proceedings of the 24th International Society for Mu...

  8. [8]

    https://doi.org/10.5281/ZENODO.10265357 Nápoles López, N

    , 597–604. https://doi.org/10.5281/ZENODO.10265357 Nápoles López, N. (2017). Joseph Haydn–String Quartets Op.20–Harmonic Analysis Annotations Dataset [Dataset]. Zenodo. https://doi.org/10.5281/ZENODO.1095630 Nápoles López, N. (2022). Automatic Roman Numeral Analysis in Symbolic Music Representations [Doctoral Dissertation]. McGill University. Nápoles Lópe...

  9. [9]

    W., & Quinn, I

    White, C. W., & Quinn, I. (2016). The Yale-Classical Archives Corpus. Empirical Musicology Review , 11 (1),

  10. [10]

    https://doi.org/10.18061/emr.v11i1.4958