Pith. sign in

REVIEW 1 major objections

A close-up comparison of the misclassification error distance and the adjusted Rand index for external clustering evaluation

T0 review · 1 major / 0 minor · reviewed 2026-05-24 · grok-4.3

Pith's one-line read Misclassification error distance and adjusted Rand index measure different aspects of clustering agreement

desk verdict This is a careful side-by-side comparison of two existing clustering metrics that clarifies their population behavior and flags some prior misconceptions, but the simulation results rest on untested design choices. read the letter →

arxiv 1907.11505 v1 pith:3PYXOWMP submitted 2019-07-26 stat.ML cs.LG

classification stat.MLcs.LG
keywords clusteringevaluationmisclassificationerrordistanceadjustedRandindexexternalvalidationsimulationstudypopulationorigins
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper compares the misclassification error distance and the adjusted Rand index, two common criteria for evaluating clustering algorithms. It traces both back to their population origins and examines their properties through data analysis examples and particular cases in detail. An exhaustive simulation study inspects the criteria distributions and reveals some previous misconceptions about how they behave.

What carries the argument

Population origins of the misclassification error distance and the adjusted Rand index, which underpin their finite-sample versions and differing behaviors

What would settle it

A concrete dataset or parameter regime in which the two criteria always produce identical rankings of candidate clusterings, contrary to the disagreements found in the paper's examples and simulations.

Watch

Extended reading notes

Core claim

Starting from their population origins, the misclassification error distance and the adjusted Rand index are shown to have distinct properties and to produce different conclusions about clustering performance in specific cases, with simulations exposing prior misconceptions about their equivalence or relative merits.

Load-bearing premise

The chosen simulation distributions and parameter ranges are representative enough to expose genuine misconceptions rather than artifacts of the simulation design.

Editorial extensions

If this is right

  • Selecting one criterion over the other can change which clustering solution is preferred for the same data
  • Some previously reported advantages or behaviors of either measure do not hold under the detailed distributional analysis
  • Particular cases exist where the two measures reach opposite conclusions about clustering quality

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same population-origin comparison could be applied to additional external clustering indices
  • Reporting results under both criteria would give a more complete picture of clustering performance in applications
  • The simulation design could be extended to compare internal validation measures in a similar way
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 0 minor

Summary. The paper compares the misclassification error distance and the adjusted Rand index for external clustering evaluation. It derives their population versions, presents multiple data analysis examples and particular cases in detail, and conducts an exhaustive simulation study to inspect the distributions of the criteria and identify previous misconceptions about their properties and differences.

Significance. If the simulation results hold under representative designs, the work supplies a clearer understanding of what these two widely used metrics actually measure, their behaviors under varying conditions, and corrections to prior misconceptions. The population-level derivations and exhaustive simulation approach constitute a strength, providing a systematic rather than ad-hoc comparison.

major comments (1)
  1. [Simulation study] Simulation study section: the central claim that the exhaustive simulations reveal intrinsic misconceptions rests on the chosen generative models and parameter ranges being representative. No sensitivity analysis or explicit coverage argument is provided to rule out design artifacts (e.g., limited overlap, balanced sizes, or low dimensionality), which directly affects whether the reported discrepancies are general or conditional on the simulation setup.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the constructive feedback on our manuscript comparing the misclassification error distance and the adjusted Rand index. We address the single major comment below.

read point-by-point responses
  1. Referee: [Simulation study] Simulation study section: the central claim that the exhaustive simulations reveal intrinsic misconceptions rests on the chosen generative models and parameter ranges being representative. No sensitivity analysis or explicit coverage argument is provided to rule out design artifacts (e.g., limited overlap, balanced sizes, or low dimensionality), which directly affects whether the reported discrepancies are general or conditional on the simulation setup.

    Authors: We agree that an explicit argument for the representativeness of the simulation design would strengthen the generalizability of the reported discrepancies. The generative models and parameter ranges were chosen to align with standard setups in the clustering evaluation literature (varying cluster numbers, overlap levels, and sample sizes), but we did not provide a dedicated sensitivity analysis or coverage discussion. In revision we will add a subsection justifying the design choices with references to prior work and include targeted additional simulations that vary dimensionality and cluster balance to verify that the key differences between the two criteria remain consistent. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: analysis of two pre-existing metrics via population versions and simulations

full rationale

The paper compares the misclassification error distance and adjusted Rand index, both established external clustering criteria. It derives their population versions from standard definitions, presents data examples, and runs simulations to inspect distributions. No step reduces a claimed prediction or result to a fitted parameter or self-citation defined inside the paper; the work is self-contained against external benchmarks and does not invoke load-bearing self-citations or ansatzes. The reader's assessment of score 1.0 is consistent with this finding of at most minor non-load-bearing elements.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The paper rests on standard statistical assumptions about clustering metrics and simulation design rather than introducing new free parameters or entities.

assumptions (1)
  • domain assumption Population-level definitions of both metrics exist and can be derived in closed form.
    The investigation starts from population origins.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A close-up comparison of the misclassification error distance and the adjusted Rand index for external clustering evaluation." pith.science (2026). https://pith.science/paper/3PYXOWMP

@misc{pith2026190711505,
  author       = {Pith},
  title        = {Pith review of: A close-up comparison of the misclassification error distance and the adjusted Rand index for external clustering evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3PYXOWMP}},
  note         = {Machine review of arXiv:1907.11505}
}
read the original abstract

The misclassification error distance and the adjusted Rand index are two of the most commonly used criteria to evaluate the performance of clustering algorithms. This paper provides an in-depth comparison of the two criteria, aimed to better understand exactly what they measure, their properties and their differences. Starting from their population origins, the investigation includes many data analysis examples and the study of particular cases in great detail. An exhaustive simulation study allows inspecting the criteria distributions and reveals some previous misconceptions.

Discussion (0). Sign in to comment.

Pith tools

Reviewed May 24, 2026 · model on record in the stance chip above.