Pith. sign in

REVIEW 2 cited by

Performance Impact Caused by Hidden Bias of Training Data for Recognizing Textual Entailment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.08117 v1 pith:IKZK4GN5 submitted 2018-04-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords hypothesisbiascorpusentailmenthiddentextualnullcaused
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The quality of training data is one of the crucial problems when a learning-centered approach is employed. This paper proposes a new method to investigate the quality of a large corpus designed for the recognizing textual entailment (RTE) task. The proposed method, which is inspired by a statistical hypothesis test, consists of two phases: the first phase is to introduce the predictability of textual entailment labels as a null hypothesis which is extremely unacceptable if a target corpus has no hidden bias, and the second phase is to test the null hypothesis using a Naive Bayes model. The experimental result of the Stanford Natural Language Inference (SNLI) corpus does not reject the null hypothesis. Therefore, it indicates that the SNLI corpus has a hidden bias which allows prediction of textual entailment labels from hypothesis sentences even if no context information is given by a premise sentence. This paper also presents the performance impact of NN models for RTE caused by this hidden bias.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Abductive Commonsense Reasoning

    cs.CL 2019-08 accept novelty 7.0 of 10

    A new benchmark, ART, shows pre-trained language models are far behind humans at choosing and generating plausible explanations of everyday narrative observations.

  2. Digital Gatekeepers: Exploring Large Language Model's Role in Immigration Decisions

    cs.CL 2025-06 conditional novelty 6.0 of 10

    GPT-3.5 and GPT-4 approximate human immigration preferences in a discrete choice experiment but exhibit systematic biases toward privileged nationalities and occupations.

Pith tools