Pith. sign in

REVIEW 2 cited by

Stress Test Evaluation for Natural Language Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.00692 v3 pith:C3SQDOQF submitted 2018-06-02 cs.CL

classification cs.CL
keywords languagemodelsnaturalevaluationstressinferencetasktests
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Natural language inference (NLI) is the task of determining if a natural language hypothesis can be inferred from a given premise in a justifiable manner. NLI was proposed as a benchmark task for natural language understanding. Existing models perform well at standard datasets for NLI, achieving impressive results across different genres of text. However, the extent to which these models understand the semantic content of sentences is unclear. In this work, we propose an evaluation methodology consisting of automatically constructed "stress tests" that allow us to examine whether systems have the ability to make real inferential decisions. Our evaluation of six sentence-encoder models on these stress tests reveals strengths and weaknesses of these models with respect to challenging linguistic phenomena, and suggests important directions for future work in this area.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Linguistics and Human Brain: A Perspective of Computational Neuroscience

    q-bio.NC 2026-02 unverdicted novelty 2.0 of 10

    A narrative review arguing that computational neuroscience, powered by LLM-based model–brain alignment, serves as the bridge between linguistic theory and neural data.

  2. From Superficial Patterns to Semantic Understanding: Fine-Tuning Language Models on Contrast Sets

    cs.CL 2025-01 conditional novelty 2.0 of 10

    Fine-tuning ELECTRA-small on a 20% subsample of a contrast set raises held-out contrast set accuracy from 74.9% to 90.7% without hurting SNLI accuracy.

Pith tools