Evaluating Classifiers Without Expert Labels

Hyun Joon Jung; Matthew Lease

arxiv: 1212.0960 · v1 · pith:HA6UWNQKnew · submitted 2012-12-05 · 💻 cs.LG · cs.IR· stat.ML

Evaluating Classifiers Without Expert Labels

Hyun Joon Jung , Matthew Lease This is my paper

classification 💻 cs.LG cs.IRstat.ML

keywords labelsclassifiersevaluationexpertlabelclassifierperformancescore

0 comments

read the original abstract

This paper considers the challenge of evaluating a set of classifiers, as done in shared task evaluations like the KDD Cup or NIST TREC, without expert labels. While expert labels provide the traditional cornerstone for evaluating statistical learners, limited or expensive access to experts represents a practical bottleneck. Instead, we seek methodology for estimating performance of the classifiers which is more scalable than expert labeling yet preserves high correlation with evaluation based on expert labels. We consider both: 1) using only labels automatically generated by the classifiers (blind evaluation); and 2) using labels obtained via crowdsourcing. While crowdsourcing methods are lauded for scalability, using such data for evaluation raises serious concerns given the prevalence of label noise. In regard to blind evaluation, two broad strategies are investigated: combine & score and score & combine methods infer a single pseudo-gold label set by aggregating classifier labels; classifiers are then evaluated based on this single pseudo-gold label set. On the other hand, score & combine methods: 1) sample multiple label sets from classifier outputs, 2) evaluate classifiers on each label set, and 3) average classifier performance across label sets. When additional crowd labels are also collected, we investigate two alternative avenues for exploiting them: 1) direct evaluation of classifiers; or 2) supervision of combine & score methods. To assess generality of our techniques, classifier performance is measured using four common classification metrics, with statistical significance tests. Finally, we measure both score and rank correlations between estimated classifier performance vs. actual performance according to expert judgments. Rigorous evaluation of classifiers from the TREC 2011 Crowdsourcing Track shows reliable evaluation can be achieved without reliance on expert labels.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Negative Ontology of True Target for Machine Learning: Towards Evaluation and Learning under Democratic Supervision
cs.LG 2026-04 unverdicted novelty 6.0

By adopting a negative ontology where the true target does not objectively exist, the paper defines Democratic Supervision and derives the EL-MIATTs framework for ML evaluation and learning with Multiple Inaccurate Tr...
Negative Ontology of True Target for Machine Learning: Towards Evaluation and Learning under Democratic Supervision
cs.LG 2026-04 unverdicted novelty 5.0

The paper posits that the true target does not exist and introduces the EL-MIATTs framework for evaluation and learning under Democratic Supervision in machine learning.
Negative Ontology of True Target for Machine Learning: Towards Evaluation and Learning under Democratic Supervision
cs.LG 2026-04 unverdicted novelty 4.0

The true target does not objectively exist in ML, so models should use multiple inaccurate true targets under democratic supervision via the EL-MIATTs framework for evaluation and learning.
LAF-Based Evaluation and UTTL-Based Learning Strategies with MIATTs
cs.LG 2026-04 unverdicted novelty 3.0

The EL-MIATTs framework offers LAF-grounded evaluation and UTTL-grounded learning strategies to handle multiple inaccurate true targets in machine learning under uncertain supervision.