REVIEW 7 cited by
Good practices for evaluation of machine learning systems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Good practices for evaluation of machine learning systems
read the original abstract
Many development decisions affect the results obtained from ML experiments: training data, features, model architecture, hyperparameters, test data, etc. Among these aspects, arguably the most important design decisions are those that involve the evaluation procedure. This procedure is what determines whether the conclusions drawn from the experiments will or will not generalize to unseen data and whether they will be relevant to the application of interest. If the data is incorrectly selected, the wrong metric is chosen for evaluation or the significance of the comparisons between models is overestimated, conclusions may be misleading or result in suboptimal development decisions. To avoid such problems, the evaluation protocol should be very carefully designed before experimentation starts. In this work we discuss the main aspects involved in the design of the evaluation protocol: data selection, metric selection, and statistical significance. This document is not meant to be an exhaustive tutorial on each of these aspects. Instead, the goal is to explain the main guidelines that should be followed in each case. We include examples taken from the speech processing field, and provide a list of common mistakes related to each aspect.
Forward citations
Cited by 7 Pith papers
-
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
L2-Bench introduces a practitioner-validated taxonomy and 1,000-task rubric benchmark showing frontier LLMs score ~85% on L2 learning-design tasks but drop to ~70% on hard items.
-
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
L2-Bench provides a 1,000-task, expert-validated rubric benchmark showing that frontier LLMs score 50–86% on applied second-language learning-design competencies, with notable weakness on open-ended tasks.
-
Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results
State-of-the-art speech recognizers match human word-error rates on Dutch child speech and outperform native listeners on older-adults and Flemish-teenager speech in a 120-utterance pilot.
-
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
EvalCards is a composable reporting schema and monitoring tool for AI evaluations, derived from 52 papers and 10 interviews, and applied to 5,816 models and 101,843 results to surface reporting gaps.
-
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
On real noisy Dutch semi-spontaneous speech, five of eight SOTA ASR models reach WER under 22%, yet five single-channel SE methods fail to improve and often degrade ASR.
-
From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed Attackers
Task-specific digit-recognition attacks show that temporal smoothing, resampling, and shredding protect spoken digits less—and differently—than read-speech word-error rates suggest.
-
Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study
Case study finds that fine-tuned ASR models outperform human listeners on Dutch dysarthric continuous speech from one speaker, lowering WER from over 70% to over 23%.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.