The Pile is a newly constructed 825 GiB dataset from 22 diverse sources that enables language models to achieve better performance on academic, professional, and cross-domain tasks than models trained on Common Crawl variants.
Journal of machine Learning research , volume=
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
A scale-free gain rule applied to variational ELBO paths recovers true factor dimensionality in partially exploratory factor analysis where raw information criteria over-factor.
Biased pay beliefs distort amenity-pay tradeoffs in stated preferences, with full disclosure restoring alignment to full-info benchmarks while short-term info does not.
A new LLM-based system generates and validates runnable modular MCMC samplers directly from natural-language Bayesian model descriptions, reporting success on 120 of 132 benchmark models.
BoolXLLM augments an existing Boolean rule learner with LLMs for feature selection, discretization thresholds, and natural-language rule translation to improve interpretability while preserving accuracy.
Mixtures of convolutional measures on low-dimensional affine spaces admit unique identifiability in semi-parametric settings and posterior contraction rates under convex polytope support assumptions in a well-specified Bayesian regime.
ALLaVA creates 1.3M GPT4V-synthesized samples enabling 4B VLMs to achieve competitive results on 17 benchmarks and match 7B/13B models on some tasks.
Embeddings reliably capture authorial stylistic features in French literary texts, and these signals persist after LLM rewriting while showing model-specific patterns.
Graph-augmented LLMs using a political knowledge graph improve ideology prediction accuracy for Swiss MPs by incorporating relational data beyond text alone.
Large-scale computational comparison of two major Holocaust oral history collections shows both expected differences and significant overlaps in interview structure, yielding a replicable framework for archive analysis.
citing papers explorer
-
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
The Pile is a newly constructed 825 GiB dataset from 22 diverse sources that enables language models to achieve better performance on academic, professional, and cross-domain tasks than models trained on Common Crawl variants.
-
Recovering Latent Structures after Variational Bayesian Variable Selection: Fit Assessment and Factor-Number Selection in Partially Exploratory Factor Analysis
A scale-free gain rule applied to variational ELBO paths recovers true factor dimensionality in partially exploratory factor analysis where raw information criteria over-factor.
-
Pay Beliefs and the Amenity-Pay Tradeoff
Biased pay beliefs distort amenity-pay tradeoffs in stated preferences, with full disclosure restoring alignment to full-info benchmarks while short-term info does not.
-
AI4BayesCode: From Natural Language Descriptions to Validated Modular Stateful Bayesian Samplers
A new LLM-based system generates and validates runnable modular MCMC samplers directly from natural-language Bayesian model descriptions, reporting success on 120 of 132 benchmark models.
-
BoolXLLM: LLM-Assisted Explainability for Boolean Models
BoolXLLM augments an existing Boolean rule learner with LLMs for feature selection, discretization thresholds, and natural-language rule translation to improve interpretability while preserving accuracy.
-
Learning Mixtures of Nonparametric and Convolutional Measures on Effectively Low-dimensional Affine Spaces
Mixtures of convolutional measures on low-dimensional affine spaces admit unique identifiability in semi-parametric settings and posterior contraction rates under convex polytope support assumptions in a well-specified Bayesian regime.
-
ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
ALLaVA creates 1.3M GPT4V-synthesized samples enabling 4B VLMs to achieve competitive results on 17 benchmarks and match 7B/13B models on some tasks.
-
Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings
Embeddings reliably capture authorial stylistic features in French literary texts, and these signals persist after LLM rewriting while showing model-specific patterns.
-
Graph-Augmented LLMs for Swiss MP Ideology Prediction
Graph-augmented LLMs using a political knowledge graph improve ideology prediction accuracy for Swiss MPs by incorporating relational data beyond text alone.
-
The Shape of Testimony: A Scalable Framework for Oral History Archive Comparison
Large-scale computational comparison of two major Holocaust oral history collections shows both expected differences and significant overlaps in interview structure, yielding a replicable framework for archive analysis.