Pith. sign in

REVIEW 1 cited by

No News is Good News: A Critique of the One Billion Word Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.12609 v1 pith:WNB7C3N6 submitted 2021-10-25 cs.CL cs.LG

classification cs.CLcs.LG
keywords languagenewsabilitybenchmarkbillioncrawlmodelingmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The One Billion Word Benchmark is a dataset derived from the WMT 2011 News Crawl, commonly used to measure language modeling ability in natural language processing. We train models solely on Common Crawl web scrapes partitioned by year, and demonstrate that they perform worse on this task over time due to distributional shift. Analysis of this corpus reveals that it contains several examples of harmful text, as well as outdated references to current events. We suggest that the temporal nature of news and its distribution shift over time makes it poorly suited for measuring language modeling ability, and discuss potential impact and considerations for researchers building language models and evaluation datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decomposed Entailment for Factuality Checking and Hallucination Detection

    cs.CL 2026-08 conditional novelty 6.0 of 10

    HallDetect detects source-grounded hallucinations by decomposing responses into atomic claims and verifying each with a compact NLI model over multi-scale source chunks, outperforming frugal generative baselines on th...

Pith tools