{"id":"eccfbab0-8384-4f63-bb3a-29ad6bcf0b58","arxiv_id":"2607.04523","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Transformer LMs track property prevalence, but only GPT-4 recovers the principled-vs-statistical generic distinction after controlling for prevalence.","lead":"Language models detect how common a property is for a category, but only GPT-4 also recovers the human distinction between properties that are true in principle versus merely statistically. This suggests large-scale language statistics can induce sophisticated causal-like conceptual structure.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"GPT-4 residual success may reflect residual prevalence/cue-validity confounds or prompt compliance rather than induction of causal structure.","rationale":"The reader correctly isolates the probe-validity assumption as the weakest link. My concern sharpens the same point: even granting that the integer-rating prompt is a reasonable operationalization, the paper’s own human analyses (Fig. 1B) show that prevalence alone is insufficient; cue-validity must also be controlled. The GPT-4 residual is reported only after prevalence control. Without the joint residual (or the item list that would allow it), the leap from residual association to “sophisticated causal models … from language alone” (§5) remains conditional. No stronger internal inconsistency or red-flag is present; the human replication is solid and the smaller-model failures are cleanly shown. Hence the verdict stays CONDITIONAL and I agree with the reader’s diagnosis.","tokens_in":6882,"tokens_out":534,"duration_ms":5469,"concrete_test":"Release the 208 statements + raw GPT-4 ratings (15 samples each). Re-run the exact Fig. 3 regressions after residualizing model truth ratings on both human prevalence and cue-validity simultaneously; if the residual property-type coefficient falls below t≈2 or item-level r with human by-virtue drops below ~.4, the claim that GPT-4 has induced the distinction beyond statistical confounds weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that GPT-4 succeeds at the principled-vs-statistical distinction (abstract; §4 residual t=5.18, p<.00001; item r=.61) rests on residual association of model by-virtue ratings with property type after controlling only for human prevalence ratings (Fig. 3). Humans also show residual by-virtue prediction after prevalence control (Fig. 1B purple bar), but the paper never reports the analogous residual for GPT-4 after jointly residualizing both human prevalence and cue-validity (the two confounds Prasada et al. 2013 emphasize and that the human design collected). Because GPT-4 is closed and the 208-item list is unreleased, it is impossible to verify whether residual co-occurrence statistics or RLHF-style prompt compliance, rather than a causal model, drive the residual. The smaller models’ cosine probes already collapse once prevalence is partialled (Fig. 2 purple bars), so the leap to “sophisticated causal models” for GPT-4 is under-supported by the reported controls.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper asks whether the principled-vs-statistical distinction for generic statements (e.g., “airplanes have wings” vs. “airplanes have passengers”) can be recovered from language statistics alone. Human ratings of 208 generics replicate Prasada et al. (2013): by-virtue-of truth judgments continue to predict property type after prevalence is controlled (Fig. 1B). Cosine similarities from decontextualized embeddings of BERT-family and GPT-2 models track prevalence but lose the residual property-type association once prevalence is partialled (Fig. 2). Direct integer ratings from GPT-3.5 show a marginal residual; GPT-4 retains a robust residual (t=5.18) and item-level correlation r=.61 with human by-virtue ratings (Fig. 3, §4). The authors conclude that sufficiently large LMs can induce the distinction and, by extension, sophisticated causal models from language.","tokens_in":7140,"tokens_out":976,"duration_ms":18490,"significance":"If the residual GPT-4 result is robust, the work supplies an existence proof that an associative system trained only on text can recover a conceptual distinction previously argued to be unlearnable (and perhaps unrepresentable) by association. That result would be of direct interest to distributional semantics, language evolution, and debates about whether next-token prediction induces world models. The human replication is clean, the regression design is transparent, and the item-level correlations are reported. These strengths make the paper worth publishing once the load-bearing controls and operationalizations are tightened.","major_comments":[{"comment":"§3.2 / Fig. 3: The central claim that GPT-4 “succeeds” rests on residual association of model by-virtue ratings with property type after controlling only for human prevalence. The human design also collected cue-validity ratings, and Prasada et al. (2013) treat both prevalence and cue-validity as confounds. The paper never reports the analogous residual for GPT-4 after jointly residualizing prevalence and cue-validity. Without that control the residual could still be driven by residual co-occurrence statistics rather than causal structure.","section":null},{"comment":"§3.2 Methods: Cosine similarity is obtained after removing the top k=7 principal components (“all-but-the-top”). k is a free parameter chosen without sensitivity analysis or justification that the residual geometry still indexes the same conceptual distinction humans make with by-virtue judgments. Because the smaller models already collapse once prevalence is partialled, any claim that the GPT-4 residual reflects induction of causal models rather than residual co-occurrence or prompt compliance requires showing that the result is stable across reasonable k (or an alternative decontextualization).","section":null},{"comment":"§4 / General Discussion: The leap from residual association to “sophisticated causal models of item-property relations” is under-supported by the reported evidence. The paper shows that GPT-4 ratings continue to predict property type after prevalence control; it does not show that the model represents the generative type-token or causal structure that Prasada and colleagues attribute to humans. A more cautious interpretation (residual statistical sensitivity that survives prevalence control) would still be interesting and would better match the data.","section":null}],"minor_comments":[{"comment":"Fig. 2 caption: “analogous models used in Fig. 2” is self-referential; should point to Fig. 1B.","section":null},{"comment":"Table 1: training-corpus sizes for GPT-3.5/4 are listed as “Unknown”; a brief note on why OpenAI API models cannot be compared on the same footing as the open models would help readers.","section":null},{"comment":"§2.3: the U-shaped coefficient pattern is described clearly, but the exact regression formulas (or a short methods appendix) would make the model comparisons fully reproducible.","section":null},{"comment":"References: several arXiv preprints are cited without final venue or year; update where possible.","section":null}],"recommendation":"major_revision","confidential_remarks":"The 208-item list and exact prompts are not released; for a closed model like GPT-4 this makes independent verification difficult. I would encourage the authors (or the editor) to require deposition of the item set and the raw model ratings as a condition of acceptance. The paper is a good fit for a cognitive-science or computational-linguistics venue that values conceptual distinctions; it is less of a pure systems paper."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new result is simple and useful: after partialling human prevalence, only GPT-4 keeps a clear residual link between its by-virtue ratings and the principled/statistical labels (t=5.18; item r=.61). Smaller models and cosine probes collapse once prevalence is controlled. That scale pattern is the paper's contribution.\n\nWhat they do well is the human side. They replicate Prasada et al. cleanly (Fig. 1), collect the right controls (prevalence and cue-validity), and show the residual human by-virtue effect survives prevalence. The model regressions are transparent and the residual analysis is the right test. The claim that language statistics can, in principle, support the distinction is therefore on firmer ground than the older 'unlearnable by association' rhetoric allowed.\n\nSoft spots are real but not fatal. The all-but-the-top decontextualization (k=7) is a free choice with no sensitivity check, and the GPT prompts are closed-model black boxes. More importantly, they never residualize GPT-4 jointly on both prevalence and cue-validity—the two confounds the human design itself collected. So the residual could still be leftover co-occurrence or prompt compliance rather than a causal model. The interpretive language about 'sophisticated causal models' and 'bootstrapping core conceptual distinctions' runs ahead of the controls. Item list and code are also unreleased, which blocks easy re-analysis.\n\nNone of that erases the residual GPT-4 pattern or the clean human data. The paper is for people who care about what distributional models can recover about conceptual structure, and for anyone still arguing that certain distinctions are definitionally unlearnable from language. It is short, readable, and the core stats are inspectable.\n\nI would send it to peer review. The residual finding is new enough and the design careful enough to deserve referee time; the overclaim and missing joint control are fixable. Worth reading and citing for the empirical pattern, with the usual caveats about what residual association actually means.","headline":"Clean human replication plus residual GPT-4 signal after prevalence control; the leap to 'causal models' is under-controlled but the empirical pattern is real and worth engaging.","tokens_in":7733,"tokens_out":518,"would_cite":true,"duration_ms":6264,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Only GPT-4 recovers the principled-versus-statistical distinction from language once prevalence is controlled.","keywords":["generics","principled properties","statistical properties","distributional semantics","language models","world models","prevalence"],"falsifier":"A controlled comparison showing that GPT-4's residual association with property type disappears once prevalence, cue validity, and surface co-occurrence are jointly partialled, or that an equivalently large model trained on scrambled co-occurrence statistics still produces the same residual signal.","tokens_in":7746,"feed_emoji":"📚","tokens_out":833,"duration_ms":7198,"temperature":0.7,"pith_summary":"People treat some generics as true by virtue of category membership (airplanes have wings) and others as merely statistical (airplanes have passengers). The distinction has been argued to be unlearnable from language structure alone. This paper tests whether distributional language models can recover it. All tested models track how common an item-property pair is, yet after prevalence is partialled out the residual principled-versus-statistical signal is weak or absent until GPT-4, whose by-virtue ratings continue to predict property type and correlate with human judgments at r = .61. The result is offered as an in-principle demonstration that sophisticated causal structure can be induced from linguistic statistics at sufficient scale, opening the possibility that language experience helps bootstrap core conceptual distinctions.","feed_headline":"GPT-4 learns principled vs statistical properties from language","feed_subtitle":"Smaller models only track prevalence; residual causal structure appears at GPT-4 scale","key_machinery":"The residual association between model truth scores (or cosine similarity after all-but-the-top decontextualization) and property type after human prevalence is partialled out; this residual is the paper's operational test of whether a model has induced something beyond co-occurrence frequency.","core_discovery":"Language models are sensitive to the statistical prevalence of item-property pairs, but the residual ability to distinguish principled from statistical generics once prevalence is controlled appears reliably only in GPT-4. Cosine-similarity probes of smaller transformers lose the distinction after controls; GPT-4's direct truth ratings retain a strong residual association (t = 5.18) and item-level correlation with human by-virtue judgments of .61.","pith_inferences":["The jump from GPT-3.5 to GPT-4 suggests a sharp threshold rather than smooth scaling; identifying the precise architectural or data change that produces residual structure would clarify how causal models form.","If language can induce the distinction, cross-linguistic corpora differing in generic density or morphological marking of kinds should produce measurable differences in model residual associations.","The same residual-control method could be applied to other allegedly unlearnable conceptual distinctions (essentialism, teleology) to test how far linguistic statistics reach."],"forward_implications":["If the residual GPT-4 signal reflects genuine causal structure, next-token prediction at scale can induce world models that separate principled from statistical category properties.","Languages may be structured so that distributional statistics alone can bootstrap the distinction, making language a source of conceptual architecture rather than only of generic facts.","Model scale and training regime become theoretically relevant variables for when sophisticated conceptual distinctions emerge from text alone.","Human conceptual development may receive more scaffolding from linguistic input than nativist accounts of the principled-statistical distinction have allowed."],"fun_headline_variants":["Only GPT-4 learns principled vs statistical distinction from language","GPT-4 retains residual principled signal after prevalence controls","Smaller LMs track prevalence; GPT-4 alone captures by-virtue structure","Models fail principled-statistical test until GPT-4 scale","GPT-4 alone separates principled generics from statistical ones"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That cosine similarity of decontextualized embeddings and integer rating prompts to GPT models measure the same principled-versus-statistical distinction people express with by-virtue-of judgments, rather than residual co-occurrence or prompt-following.","fun_headline_variants_meta":{"raw":{"variants":["Only GPT-4 learns principled vs statistical distinction from language","GPT-4 retains residual principled signal after prevalence controls","Smaller LMs track prevalence; GPT-4 alone captures by-virtue structure","Models fail principled-statistical test until GPT-4 scale","GPT-4 alone separates principled generics from statistical ones"]},"model":"grok-4.5","effort":"low","cost_usd":0.00479,"raw_usage":{"total_tokens":1341,"prompt_tokens":718,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":47900000,"prompt_tokens_details":{"text_tokens":718,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":551,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":718,"tokens_out":72,"duration_ms":6312,"temperature":1.0,"reasoning_tokens":551,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T18:01:38.605052+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A controlled comparison showing that GPT-4's residual association with property type disappears once prevalence, cue validity, and surface co-occurrence are jointly partialled, or that an equivalently large model trained on scrambled co-occurrence statistics still produces the same residual signal.","supporting_citations":[],"review_version":1}