Pith. sign in

REVIEW 3 cited by

An Approach to Multiple Comparison Benchmark Evaluations that is Stable Under Manipulation of the Comparate Set

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.11921 v1 pith:VE6J43XA submitted 2023-05-19 stat.ME cs.AIcs.LGcs.PF

classification stat.MEcs.AIcs.LGcs.PF
keywords resultscomparisonmultiplebenchmarkcomparisonsalgorithmsapproachapproaches
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The measurement of progress using benchmarks evaluations is ubiquitous in computer science and machine learning. However, common approaches to analyzing and presenting the results of benchmark comparisons of multiple algorithms over multiple datasets, such as the critical difference diagram introduced by Dem\v{s}ar (2006), have important shortcomings and, we show, are open to both inadvertent and intentional manipulation. To address these issues, we propose a new approach to presenting the results of benchmark comparisons, the Multiple Comparison Matrix (MCM), that prioritizes pairwise comparisons and precludes the means of manipulating experimental results in existing approaches. MCM can be used to show the results of an all-pairs comparison, or to show the results of a comparison between one or more selected algorithms and the state of the art. MCM is implemented in Python and is publicly available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Statistical comparisons of time-series feature sets on classification tasks

    stat.ME 2026-08 conditional novelty 6.0 of 10

    Across 124 time-series classification datasets, six open-source feature sets perform mostly equivalently, with tsfresh winning most often and simple quantile/FFT baselines competitive on several problems.

  2. Scaling Time Series Classification via XAI-Driven Data Reduction

    cs.LG 2026-07 conditional novelty 6.0 of 10

    drXAI uses XAI attributions from a fast classifier to choose important channels/time points, achieving 80–90% data reduction with comparable classification accuracy.

  3. A Standardized Benchmark for Skeleton-Based Rehabilitation Assessment Using Deep Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Rehab-Pile, a standardized archive of 60 rehabilitation skeleton datasets with reproducible baselines, reports that the authors' lightweight LITEMV model wins on efficiency and usually on accuracy.

Pith tools