REVIEW 1 cited by
No News is Good News: A Critique of the One Billion Word Benchmark
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The One Billion Word Benchmark is a dataset derived from the WMT 2011 News Crawl, commonly used to measure language modeling ability in natural language processing. We train models solely on Common Crawl web scrapes partitioned by year, and demonstrate that they perform worse on this task over time due to distributional shift. Analysis of this corpus reveals that it contains several examples of harmful text, as well as outdated references to current events. We suggest that the temporal nature of news and its distribution shift over time makes it poorly suited for measuring language modeling ability, and discuss potential impact and considerations for researchers building language models and evaluation datasets.
Forward citations
Cited by 1 Pith paper
-
Decomposed Entailment for Factuality Checking and Hallucination Detection
HallDetect detects source-grounded hallucinations by decomposing responses into atomic claims and verifying each with a compact NLI model over multi-scale source chunks, outperforming frugal generative baselines on th...
Discussion (0). Continue with ORCID to comment.