Pith. sign in

REVIEW 7 cited by

Good practices for evaluation of machine learning systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.03700 v1 pith:NNNGUHXP submitted 2024-12-04 cs.LG

Good practices for evaluation of machine learning systems

classification cs.LG
keywords dataevaluationaspectsdecisionswillconclusionsdesigndevelopment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Many development decisions affect the results obtained from ML experiments: training data, features, model architecture, hyperparameters, test data, etc. Among these aspects, arguably the most important design decisions are those that involve the evaluation procedure. This procedure is what determines whether the conclusions drawn from the experiments will or will not generalize to unseen data and whether they will be relevant to the application of interest. If the data is incorrectly selected, the wrong metric is chosen for evaluation or the significance of the comparisons between models is overestimated, conclusions may be misleading or result in suboptimal development decisions. To avoid such problems, the evaluation protocol should be very carefully designed before experimentation starts. In this work we discuss the main aspects involved in the design of the evaluation protocol: data selection, metric selection, and statistical significance. This document is not meant to be an exhaustive tutorial on each of these aspects. Instead, the goal is to explain the main guidelines that should be followed in each case. We include examples taken from the speech processing field, and provide a list of common mistakes related to each aspect.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education

    cs.CY 2026-07 conditional novelty 7.0

    L2-Bench introduces a practitioner-validated taxonomy and 1,000-task rubric benchmark showing frontier LLMs score ~85% on L2 learning-design tasks but drop to ~70% on hard items.

  2. L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education

    cs.CY 2026-07 conditional novelty 7.0

    L2-Bench provides a 1,000-task, expert-validated rubric benchmark showing that frontier LLMs score 50–86% on applied second-language learning-design competencies, with notable weakness on open-ended tasks.

  3. Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results

    cs.CL 2026-07 conditional novelty 6.0

    State-of-the-art speech recognizers match human word-error rates on Dutch child speech and outperform native listeners on older-adults and Flemish-teenager speech in a 120-utterance pilot.

  4. Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

    cs.AI 2026-06 unverdicted novelty 6.0

    EvalCards is a composable reporting schema and monitoring tool for AI evaluations, derived from 52 papers and 10 interviews, and applied to 5,816 models and 101,843 results to surface reporting gaps.

  5. A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition

    eess.AS 2026-03 conditional novelty 5.5

    On real noisy Dutch semi-spontaneous speech, five of eight SOTA ASR models reach WER under 22%, yet five single-channel SE methods fail to improve and often degrade ASR.

  6. From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed Attackers

    eess.AS 2026-07 conditional novelty 5.0

    Task-specific digit-recognition attacks show that temporal smoothing, resampling, and shredding protect spoken digits less—and differently—than read-speech word-error rates suggest.

  7. Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study

    cs.CL 2026-06 unverdicted novelty 4.0

    Case study finds that fine-tuned ASR models outperform human listeners on Dutch dysarthric continuous speech from one speaker, lowering WER from over 70% to over 23%.