Pith. sign in

REVIEW 1 cited by

Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.03256 v2 pith:TADCGNG3 submitted 2022-10-06 cs.CL

classification cs.CL
keywords negationsuitetestlanguagesub-clausalannotationcapabilitiesconstructed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Negation is poorly captured by current language models, although the extent of this problem is not widely understood. We introduce a natural language inference (NLI) test suite to enable probing the capabilities of NLP methods, with the aim of understanding sub-clausal negation. The test suite contains premise--hypothesis pairs where the premise contains sub-clausal negation and the hypothesis is constructed by making minimal modifications to the premise in order to reflect different possible interpretations. Aside from adopting standard NLI labels, our test suite is systematically constructed under a rigorous linguistic framework. It includes annotation of negation types and constructions grounded in linguistic theory, as well as the operations used to construct hypotheses. This facilitates fine-grained analysis of model performance. We conduct experiments using pre-trained language models to demonstrate that our test suite is more challenging than existing benchmarks focused on negation, and show how our annotation supports a deeper understanding of the current NLI capabilities in terms of negation and quantification.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A survey of NLU diagnostics benchmarks finds no shared naming convention or standard set of linguistic phenomena, and asks whether the field should build an ISO-like evaluation standard.

Pith tools