Pith. sign in

REVIEW 13 cited by

Annotation Artifacts in Natural Language Inference Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1803.02324 v2 pith:W2TAQSIR submitted 2018-03-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords inferencelanguagenaturaldatahypothesispremisealoneanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large-scale datasets for natural language inference are created by presenting crowd workers with a sentence (premise), and asking them to generate three new sentences (hypotheses) that it entails, contradicts, or is logically neutral with respect to. We show that, in a significant portion of such data, this protocol leaves clues that make it possible to identify the label by looking only at the hypothesis, without observing the premise. Specifically, we show that a simple text categorization model can correctly classify the hypothesis alone in about 67% of SNLI (Bowman et. al, 2015) and 53% of MultiNLI (Williams et. al, 2017). Our analysis reveals that specific linguistic phenomena such as negation and vagueness are highly correlated with certain inference classes. Our findings suggest that the success of natural language inference models to date has been overestimated, and that the task remains a hard open problem.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems

    cs.CR 2026-07 accept novelty 7.0 of 10

    Once attack fragments are locally benign and ε-indistinguishable from benign traffic, no local monitor can separate them (TPR−FPR≤ε); the signal reappears only in the right assembled representation.

  2. Language Models as Knowledge Bases?

    cs.CL 2019-09 accept novelty 7.0 of 10

    BERT stores relational knowledge extractable via cloze queries without fine-tuning and matches supervised baselines on open-domain QA tasks.

  3. idSCD: Identifying Training Datasets through Semantic Correlation Descriptors

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    idSCD uses semantic correlation descriptors to perform dataset membership inference by comparing learned semantic structures, outperforming baselines in NLI, emotion, and medical text experiments.

  4. The Knowledge-Reasoning Dissociation: Fundamental Limitations of LLMs in Clinical Natural Language Inference

    cs.AI 2025-08 reject novelty 6.0 of 10

    Across four clinical inference tasks, six LLMs answer paired knowledge probes at 92% accuracy but the main reasoning tasks at 25%, indicating a systematic knowledge-reasoning gap.

  5. Moment Alignment: Unifying Gradient and Hessian Matching for Domain Generalization

    cs.LG 2025-06 reject novelty 6.0 of 10

    A unified moment-alignment theory bounds target-domain error by cross-domain differences in loss derivatives, and the new CMA algorithm implements exact gradient and Hessian matching in closed form.

  6. ART: Automatic multi-step reasoning and tool-use for large language models

    cs.CL 2023-03 unverdicted novelty 6.0 of 10

    ART automatically generates multi-step reasoning programs with tool integration for LLMs, yielding substantial gains over few-shot and auto-CoT prompting on BigBench and MMLU while matching hand-crafted CoT on most tasks.

  7. Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets

    cs.CL 2019-08 conditional novelty 6.0 of 10

    Annotator identity functions as a shortcut in crowdsourced NLU benchmarks, and models often fail to generalize to examples from annotators absent from training.

  8. Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

    cs.CL 2026-07 reject novelty 5.0 of 10

    Benchmark items are heterogeneous; the authors annotate them with three LLM judges and use the labels to build filtered subsets, but the validation evidence is weak and the comparison method is flawed.

  9. PROBE: Benchmarking Code Generation in Large Language Models

    cs.SE 2026-07 conditional novelty 5.0 of 10

    A multi-language evaluation framework measuring correctness, solution proximity, and code quality finds current LLMs pass at most ~0.70 per language and worsen sharply with problem difficulty.

  10. Mass-Editing Memory with Attention in Transformers: A cross-lingual exploration of knowledge

    cs.CL 2025-02 conditional novelty 5.0 of 10

    MEMAT combines MEMIT weight edits with optimized attention-head corrections, improving cross-lingual success and magnitude metrics over MEMIT in English and Catalan.

  11. Knowledge Enhanced Attention for Robust Natural Language Inference

    cs.CL 2019-08 conditional novelty 5.0 of 10

    Injecting WordNet lexical relations as bias terms into multi-head attention improves accuracy on the adversarial SNLI test set, with BERT reaching 94.1%, equal to estimated human performance.

  12. Investigating Biases in Textual Entailment Datasets

    cs.CL 2019-06 unverdicted novelty 5.0 of 10

    Hypothesis-only classification reaches 64% accuracy on SNLI, revealing dataset biases in SNLI and MultiNLI that the authors quantify and propose a simple mitigation for.

  13. Unpacking the Resilience of SNLI Contradiction Examples to Attacks

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Adding a single universal trigger word to SNLI hypotheses crashes ELECTRA's accuracy on entailment and neutral classes, barely affects contradiction, and trigger-augmented fine-tuning recovers the lost accuracy.

Pith tools