Pith. sign in

REVIEW 2 cited by

SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.05326 v1 pith:3IRNSILW submitted 2018-08-16 cs.CL

classification cs.CL
keywords inferenceadversarialcommonsensedatasetgroundedfilteringhumanslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Given a partial description like "she opened the hood of the car," humans can reason about the situation and anticipate what might come next ("then, she examined the engine"). In this paper, we introduce the task of grounded commonsense inference, unifying natural language inference and commonsense reasoning. We present SWAG, a new dataset with 113k multiple choice questions about a rich spectrum of grounded situations. To address the recurring challenges of the annotation artifacts and human biases found in many existing datasets, we propose Adversarial Filtering (AF), a novel procedure that constructs a de-biased dataset by iteratively training an ensemble of stylistic classifiers, and using them to filter the data. To account for the aggressive adversarial filtering, we use state-of-the-art language models to massively oversample a diverse set of potential counterfactuals. Empirical results demonstrate that while humans can solve the resulting inference problems with high accuracy (88%), various competitive models struggle on our task. We provide comprehensive analysis that indicates significant opportunities for future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. OPT: Open Pre-trained Transformer Language Models

    cs.CL 2022-05 unverdicted novelty 7.0 of 10

    OPT releases open decoder-only transformers up to 175B parameters that match GPT-3 performance at one-seventh the carbon cost, along with code and training logs.

  2. LENS: Learning Ensemble Confidence from Neural States for Multi-LLM Answer Integration

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A linear probe over layer-wise logit-lens probabilities ranks LLM confidence well enough to slightly beat voting or probability-based baselines in QA ensembles.

Pith tools