Pith. sign in

REVIEW 1 cited by

MSciNLI: A Diverse Benchmark for Scientific Natural Language Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.08066 v1 pith:YVDY5JGC submitted 2024-04-11 cs.CL

classification cs.CL
keywords scientificdatasetdomainlanguagemodelsmscinlitaskdomains
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The task of scientific Natural Language Inference (NLI) involves predicting the semantic relation between two sentences extracted from research articles. This task was recently proposed along with a new dataset called SciNLI derived from papers published in the computational linguistics domain. In this paper, we aim to introduce diversity in the scientific NLI task and present MSciNLI, a dataset containing 132,320 sentence pairs extracted from five new scientific domains. The availability of multiple domains makes it possible to study domain shift for scientific NLI. We establish strong baselines on MSciNLI by fine-tuning Pre-trained Language Models (PLMs) and prompting Large Language Models (LLMs). The highest Macro F1 scores of PLM and LLM baselines are 77.21% and 51.77%, respectively, illustrating that MSciNLI is challenging for both types of models. Furthermore, we show that domain shift degrades the performance of scientific NLI models which demonstrates the diverse characteristics of different domains in our dataset. Finally, we use both scientific NLI datasets in an intermediate task transfer learning setting and show that they can improve the performance of downstream tasks in the scientific domain. We make our dataset and code available on Github.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A MISMATCHED Benchmark for Scientific Natural Language Inference

    cs.CL 2025-06 conditional novelty 5.0 of 10

    MISMATCHED is a new out-of-domain benchmark for scientific NLI spanning three non-CS domains, with best baselines at 78.17% Macro F1 and evidence that implicit-relation training helps.

Pith tools