REVIEW 3 major objections 4 minor 46 references
Failures and Successes to Learn a Core Conceptual Distinction from the Statistics of Language
T0 review · 3 major / 4 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read Only GPT-4 recovers the principled-versus-statistical distinction from language once prevalence is controlled.
desk verdict Clean human replication plus residual GPT-4 signal after prevalence control; the leap to 'causal models' is under-controlled but the empirical pattern is real and worth engaging. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The residual association between model truth scores (or cosine similarity after all-but-the-top decontextualization) and property type after human prevalence is partialled out; this residual is the paper's operational test of whether a model has induced something beyond co-occurrence frequency.
What would settle it
A controlled comparison showing that GPT-4's residual association with property type disappears once prevalence, cue validity, and surface co-occurrence are jointly partialled, or that an equivalently large model trained on scrambled co-occurrence statistics still produces the same residual signal.
Extended reading notes
Core claim
Language models are sensitive to the statistical prevalence of item-property pairs, but the residual ability to distinguish principled from statistical generics once prevalence is controlled appears reliably only in GPT-4. Cosine-similarity probes of smaller transformers lose the distinction after controls; GPT-4's direct truth ratings retain a strong residual association (t = 5.18) and item-level correlation with human by-virtue judgments of .61.
Load-bearing premise
That cosine similarity of decontextualized embeddings and integer rating prompts to GPT models measure the same principled-versus-statistical distinction people express with by-virtue-of judgments, rather than residual co-occurrence or prompt-following.
Editorial extensions
If this is right
- If the residual GPT-4 signal reflects genuine causal structure, next-token prediction at scale can induce world models that separate principled from statistical category properties.
- Languages may be structured so that distributional statistics alone can bootstrap the distinction, making language a source of conceptual architecture rather than only of generic facts.
- Model scale and training regime become theoretically relevant variables for when sophisticated conceptual distinctions emerge from text alone.
- Human conceptual development may receive more scaffolding from linguistic input than nativist accounts of the principled-statistical distinction have allowed.
Reading between the lines
- The jump from GPT-3.5 to GPT-4 suggests a sharp threshold rather than smooth scaling; identifying the precise architectural or data change that produces residual structure would clarify how causal models form.
- If language can induce the distinction, cross-linguistic corpora differing in generic density or morphological marking of kinds should produce measurable differences in model residual associations.
- The same residual-control method could be applied to other allegedly unlearnable conceptual distinctions (essentialism, teleology) to test how far linguistic statistics reach.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether the principled-vs-statistical distinction for generic statements (e.g., “airplanes have wings” vs. “airplanes have passengers”) can be recovered from language statistics alone. Human ratings of 208 generics replicate Prasada et al. (2013): by-virtue-of truth judgments continue to predict property type after prevalence is controlled (Fig. 1B). Cosine similarities from decontextualized embeddings of BERT-family and GPT-2 models track prevalence but lose the residual property-type association once prevalence is partialled (Fig. 2). Direct integer ratings from GPT-3.5 show a marginal residual; GPT-4 retains a robust residual (t=5.18) and item-level correlation r=.61 with human by-virtue ratings (Fig. 3, §4). The authors conclude that sufficiently large LMs can induce the distinction and, by extension, sophisticated causal models from language.
Significance. If the residual GPT-4 result is robust, the work supplies an existence proof that an associative system trained only on text can recover a conceptual distinction previously argued to be unlearnable (and perhaps unrepresentable) by association. That result would be of direct interest to distributional semantics, language evolution, and debates about whether next-token prediction induces world models. The human replication is clean, the regression design is transparent, and the item-level correlations are reported. These strengths make the paper worth publishing once the load-bearing controls and operationalizations are tightened.
major comments (3)
- §3.2 / Fig. 3: The central claim that GPT-4 “succeeds” rests on residual association of model by-virtue ratings with property type after controlling only for human prevalence. The human design also collected cue-validity ratings, and Prasada et al. (2013) treat both prevalence and cue-validity as confounds. The paper never reports the analogous residual for GPT-4 after jointly residualizing prevalence and cue-validity. Without that control the residual could still be driven by residual co-occurrence statistics rather than causal structure.
- §3.2 Methods: Cosine similarity is obtained after removing the top k=7 principal components (“all-but-the-top”). k is a free parameter chosen without sensitivity analysis or justification that the residual geometry still indexes the same conceptual distinction humans make with by-virtue judgments. Because the smaller models already collapse once prevalence is partialled, any claim that the GPT-4 residual reflects induction of causal models rather than residual co-occurrence or prompt compliance requires showing that the result is stable across reasonable k (or an alternative decontextualization).
- §4 / General Discussion: The leap from residual association to “sophisticated causal models of item-property relations” is under-supported by the reported evidence. The paper shows that GPT-4 ratings continue to predict property type after prevalence control; it does not show that the model represents the generative type-token or causal structure that Prasada and colleagues attribute to humans. A more cautious interpretation (residual statistical sensitivity that survives prevalence control) would still be interesting and would better match the data.
minor comments (4)
- Fig. 2 caption: “analogous models used in Fig. 2” is self-referential; should point to Fig. 1B.
- Table 1: training-corpus sizes for GPT-3.5/4 are listed as “Unknown”; a brief note on why OpenAI API models cannot be compared on the same footing as the open models would help readers.
- §2.3: the U-shaped coefficient pattern is described clearly, but the exact regression formulas (or a short methods appendix) would make the model comparisons fully reproducible.
- References: several arXiv preprints are cited without final venue or year; update where possible.
Circularity Check
No circularity: independent human labels/prevalence ratings are compared to separately probed model embeddings or ratings; residual associations are empirical, not definitional.
full rationale
The paper's central claim is an empirical residual association (property type ~ model measure | human prevalence) that appears for GPT-4 but not smaller models. Property-type labels and human prevalence/by-virtue ratings are collected from participants on a fixed 208-item corpus (replicating Prasada et al. 2013); model probes (all-but-the-top cosine similarities for BERT-family models; direct integer prompts for GPT-3.5/4) are obtained independently. No parameter is fitted to the target residual and then re-presented as a prediction; no equation equates the residual to an input by construction; no uniqueness theorem or ansatz is imported via self-citation to force the result. Citations to Prasada supply the conceptual distinction and experimental template, which the authors re-collect and re-analyze; they do not reduce the GPT-4 residual (t=5.18) to a prior claim by the same authors. The design is therefore self-contained against external benchmarks and contains no circular steps of the enumerated kinds.
Assumptions & free parameters
free parameters (3)
- k (principal components removed in all-but-the-top) =
7
- GPT rating samples per item =
15
- Human rating scale bounds =
-3 to +3
assumptions (3)
- domain assumption Principled vs. statistical generics are a real, prevalence-independent conceptual distinction recoverable from human by-virtue-of judgments (Prasada et al.).
- ad hoc to paper Cosine similarity of decontextualized transformer embeddings (or direct integer ratings from instruction-tuned GPTs) is a valid proxy for the model's representation of item-property truth and by-virtue structure.
- domain assumption Controlling for human prevalence ratings isolates the non-statistical component of the principled/statistical distinction.
Cite this review
Pith. "Pith review of Failures and Successes to Learn a Core Conceptual Distinction from the Statistics of Language." pith.science (2026). https://pith.science/paper/JPC2KOAR
@misc{pith2026260704523,
author = {Pith},
title = {Pith review of: Failures and Successes to Learn a Core Conceptual Distinction from the Statistics of Language},
year = {2026},
howpublished = {\url{https://pith.science/paper/JPC2KOAR}},
note = {Machine review of arXiv:2607.04523}
}
read the original abstract
Generic statements like "tigers are striped" and "cars have radios" communicate information that is, in general, true. However, while the first statement is true in principle, the second is true only statistically. People are exquisitely sensitive to this principled-vs-statistical distinction. It has been argued that this ability to distinguish between something being true by virtue of it being a category member versus being true because of mere statistical regularity, is a general property of people's conceptual machinery and cannot itself be learned. We investigate whether the distinction between principled and statistical properties can be learned from language itself. If so, it raises the possibility that language experience can bootstrap core conceptual distinctions and that it is possible to learn sophisticated causal models directly from language. We find that language models are all sensitive to statistical prevalence, but struggle with representing the principled-vs-statistical distinction controlling for prevalence. Until GPT-4, which succeeds.
Figures
Reference graph
Works this paper leans on
-
[1]
Bambi: A simple interface for fitting
Yarkoni, Tal and Westfall, Jake , year=. Bambi: A simple interface for fitting. OSF preprint , doi=
-
[2]
Hollander, Michelle A. and Gelman, S.A. and Raman, Lakshmi , year =. Generic Language and Judgements about Category Membership:. Language and cognitive processes , volume =. doi:10.1080/01690960802223485 , url =
-
[3]
The Physical Basis of Conceptual Representation -
Prasada, Sandeep , year =. The Physical Basis of Conceptual Representation -. Cognition , volume =. doi:10.1016/j.cognition.2021.104751 , langid =
-
[4]
Visual Cognition , pages=
Seeing colour through language: Colour knowledge in the blind and sighted , author=. Visual Cognition , pages=. 2021 , publisher=
2021
-
[5]
Reading Research Quarterly , pages=
Exposure to print and orthographic processing , author=. Reading Research Quarterly , pages=. 1989 , publisher=
1989
-
[6]
Nature , volume=
Measurement of diversity , author=. Nature , volume=. 1949 , publisher=
1949
-
[7]
arXiv preprint, arXiv:1802.01241 , year=
Semantic projection: Recovering human knowledge of multiple, distinct object features from word embeddings , author=. arXiv preprint, arXiv:1802.01241 , year=
-
[8]
Behavior Research Methods , pages=
subs2vec: Word embeddings from subtitles in 55 languages , author=. Behavior Research Methods , pages=. 2020 , publisher=
2020
Show all 46 references
-
[9]
Enriching
Bojanowski, Piotr and Grave, Edouard and Joulin, Armand and Mikolov, Tomas , year =. Enriching. arXiv:1607.04606 [cs] , eprint =
-
[10]
2018 , journal=
Learning Word Vectors for 157 Languages , author=. 2018 , journal=
2018
-
[11]
Psychological Science , volume=
Representation of colors in the blind, color-blind, and normally sighted , author=. Psychological Science , volume=. 1992 , publisher=
1992
-
[12]
Visual Cognition , volume=
Colour envisioned: Concepts of colour in the blind and sighted , author=. Visual Cognition , volume=. 2018 , publisher=
2018
-
[13]
Journal of Experimental Child Psychology , volume=
Age at onset of blindness and the development of the semantics of color names , author=. Journal of Experimental Child Psychology , volume=. 1978 , publisher=
1978
-
[14]
2019 , publisher=
De Deyne, Simon and Navarro, Danielle J and Perfors, Amy and Brysbaert, Marc and Storms, Gert , journal=. 2019 , publisher=
2019
-
[15]
1690 , address=
An essay concerning human understanding , author=. 1690 , address=
-
[16]
1740 , address=
A treatise of human nature , author=. 1740 , address=
-
[17]
Behavior Research Methods , volume=
New and updated tests of print exposure and reading abilities in college students , author=. Behavior Research Methods , volume=. 2008 , publisher=
2008
-
[18]
Wang, Alex and Cho, Kyunghyun , journal=
-
[19]
Cognition , volume=
Conceptual distinctions amongst generics , author=. Cognition , volume=. 2013 , publisher=
2013
-
[20]
Cognition , volume=
Principled and statistical connections in common sense conception , author=. Cognition , volume=. 2006 , publisher=
2006
-
[21]
2012 , month = aug, journal =
Cultural Transmission of Social Essentialism , author =. 2012 , month = aug, journal =. doi:10.1073/pnas.1208951109 , url =
2012 doi
-
[22]
Mechanisms for Thinking about Kinds, Instances of Kinds, and Kinds of Kinds , booktitle =
Prasada, Sandeep , year =. Mechanisms for Thinking about Kinds, Instances of Kinds, and Kinds of Kinds , booktitle =. doi:10.1093/acprof:oso/9780190467630.003.0012 , isbn =
-
[23]
Cognition , volume=
The development of principled connections and kind representations , author=. Cognition , volume=. 2018 , publisher=
2018
-
[24]
International Conference on Learning Representations , year=
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations , author=. International Conference on Learning Representations , year=
-
[25]
arXiv preprint arXiv:1810.04805 , year=
Bert: Pre-training of deep bidirectional transformers for language understanding , author=. arXiv preprint arXiv:1810.04805 , year=
-
[26]
Advances in neural information processing systems , volume=
Xlnet: Generalized autoregressive pretraining for language understanding , author=. Advances in neural information processing systems , volume=
-
[27]
arXiv preprint arXiv:1907.11692 , year=
Roberta: A robustly optimized bert pretraining approach , author=. arXiv preprint arXiv:1907.11692 , year=
1907 arXiv
-
[28]
arXiv preprint arXiv:1910.01108 , year=
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter , author=. arXiv preprint arXiv:1910.01108 , year=
1910 arXiv
-
[29]
OpenAI blog , volume=
Language models are unsupervised multitask learners , author=. OpenAI blog , volume=
-
[30]
2018 , journal=
Improving language understanding by generative pre-training , author=. 2018 , journal=
2018
-
[31]
2022 , url =
R: A Language and Environment for Statistical Computing , author =. 2022 , url =
2022
-
[32]
Fitting Linear Mixed-Effects Models Using
Douglas Bates and Martin M. Fitting Linear Mixed-Effects Models Using. Journal of Statistical Software , year =
-
[33]
arXiv preprint arXiv:1301.3781 , year=
Efficient estimation of word representations in vector space , author=. arXiv preprint arXiv:1301.3781 , year=
-
[34]
Proceedings of the Annual Meeting of the Cognitive Science Society , volume=
How do blind people know that blue is cold? Distributional semantics encode color-adjective associations , author=. Proceedings of the Annual Meeting of the Cognitive Science Society , volume=
-
[35]
International Conference on Learning Representations , year=
All-but-the-Top: Simple and Effective Postprocessing for Word Representations , author=. International Conference on Learning Representations , year=
-
[36]
Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=
On the Sentence Embeddings from Pre-trained Language Models , author=. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=
2020
-
[37]
Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , pages=
A large annotated corpus for learning natural language inference , author=. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , pages=
2015
-
[38]
Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , pages=
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference , author=. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , pages=
2018
-
[39]
International Conference on Learning Representations , year=
Pointer Sentinel Mixture Models , author=. International Conference on Learning Representations , year=
-
[40]
Proceedings of the 29th International Conference on Computational Linguistics , pages=
Effect of Post-processing on Contextualized Word Representations , author=. Proceedings of the 29th International Conference on Computational Linguistics , pages=
-
[41]
The Eleventh International Conference on Learning Representations , year=
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task , author=. The Eleventh International Conference on Learning Representations , year=
-
[42]
7th Annual Conference on Robot Learning , year=
Large Language Models as General Pattern Machines , author=. 7th Annual Conference on Robot Learning , year=
-
[43]
arXiv preprint arXiv:2301.08731 , year=
Can Peanuts Fall in Love with Distributional Semantics? , author=. arXiv preprint arXiv:2301.08731 , year=
-
[44]
Implicit Representations of Meaning in Neural Language Models , author=. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , pages=
-
[45]
arXiv preprint arXiv:1909.08593 , year=
Fine-tuning language models from human preferences , author=. arXiv preprint arXiv:1909.08593 , year=
1909 arXiv
-
[46]
and Angeli, Gabor and Potts, Christopher and Manning, Christopher D
Bowman, Samuel R. and Angeli, Gabor and Potts, Christopher and Manning, Christopher D. A large annotated corpus for learning natural language inference. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. doi:10.18653/v1/D15-1075
2015 doi
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.