About 983,000 public domain books from Harvard Library's Google Books digitization, containing roughly 242B tokens, were processed and released with metadata, language and topic labels, dedup hints, and post-processed OCR.
Swedigh” instead of “ Swedish
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability
About 983,000 public domain books from Harvard Library's Google Books digitization, containing roughly 242B tokens, were processed and released with metadata, language and topic labels, dedup hints, and post-processed OCR.