Pith. sign in

REVIEW 1 cited by

Negation: A Pink Elephant in the Large Language Models' Room?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.22395 v2 pith:VVN4YRB4 submitted 2025-03-28 cs.CL

classification cs.CL
keywords negationlanguagemodelsnegationsreasoningaccuracyczechdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Negations are key to determining sentence meaning, making them essential for logical reasoning. Despite their importance, negations pose a substantial challenge for large language models (LLMs) and remain underexplored. We constructed and published two new textual entailment datasets NoFEVER-ML and NoSNLI-ML in four languages (English, Czech, German, and Ukrainian) with examples differing in negation. It allows investigation of the root causes of the negation problem and its exemplification: how popular LLM model properties and language impact their inability to handle negation correctly. Contrary to previous work, we show that increasing the model size may improve the models' ability to handle negations. Furthermore, we find that both the models' reasoning accuracy and robustness to negation are language-dependent and that the length and explicitness of the premise have an impact on robustness. There is better accuracy in projective language with fixed order, such as English, than in non-projective ones, such as German or Czech. Our entailment datasets pave the way to further research for explanation and exemplification of the negation problem, minimization of LLM hallucinations, and improvement of LLM reasoning in multilingual settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding

    cs.CL 2026-01 conditional novelty 6.0 of 10

    A new Korean benchmark shows language models underperform on negated sentences, and fine-tuning on it improves their negation understanding.

Pith tools