Pith. sign in

REVIEW 3 cited by

Evaluating Models' Local Decision Boundaries via Contrast Sets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.02709 v2 pith:5NKRPZ7G submitted 2020-04-06 cs.CL

classification cs.CL
keywords setscontrastdatasettestmodelannotationdecisioncapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Standard test sets for supervised learning evaluate in-distribution generalization. Unfortunately, when a dataset has systematic gaps (e.g., annotation artifacts), these evaluations are misleading: a model can learn simple decision rules that perform well on the test set but do not capture a dataset's intended capabilities. We propose a new annotation paradigm for NLP that helps to close systematic gaps in the test data. In particular, after a dataset is constructed, we recommend that the dataset authors manually perturb the test instances in small but meaningful ways that (typically) change the gold label, creating contrast sets. Contrast sets provide a local view of a model's decision boundary, which can be used to more accurately evaluate a model's true linguistic capabilities. We demonstrate the efficacy of contrast sets by creating them for 10 diverse NLP datasets (e.g., DROP reading comprehension, UD parsing, IMDb sentiment analysis). Although our contrast sets are not explicitly adversarial, model performance is significantly lower on them than on the original test sets---up to 25\% in some cases. We release our contrast sets as new evaluation benchmarks and encourage future dataset construction efforts to follow similar annotation processes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What does AI consider praiseworthy?

    cs.CY 2024-11 conditional novelty 7.0 of 10

    LLM praise and critique of user-stated intentions are driven more by source trustworthiness than ideology, align broadly with human moral scores, and show no country-of-origin bias.

  2. The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Fixed activation probes keep near-ceiling accuracy on harmful-vs-benign corpus contrasts but fall to AUROC 0.59-0.69 on topic- and surface-matched harmful/benign pairs, so they behave as broad-risk detectors, not cont...

  3. From Superficial Patterns to Semantic Understanding: Fine-Tuning Language Models on Contrast Sets

    cs.CL 2025-01 conditional novelty 2.0 of 10

    Fine-tuning ELECTRA-small on a 20% subsample of a contrast set raises held-out contrast set accuracy from 74.9% to 90.7% without hurting SNLI accuracy.

Pith tools