REVIEW 4 major objections 5 minor 5 references
What General Intelligence Requires: Non-Reducible Constraints Across Levels of Description
T0 review · 4 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read General intelligence cannot be built by scaling or any single architectural advance, because its structural constraints are mutually non-reducible across levels of description.
desk verdict A serious, honest framework for AGI constraints, but the anti-scaling conclusion overreaches: Fodor's non-reducibility explains why levels don't reduce, not why one architecture can't jointly instantiate them. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the pairing of four evidential lenses with distinct levels of description, combined with the multiple-realisability argument from the philosophy of the special sciences. The AI-systems lens operates at the computational level, the anthropology lens at the evolutionary and developmental level, the legal lens at the normative and procedural level, and the economics lens at the incentive-theoretic level; the claim that a lens and a level are paired is the load-bearing move. Between the six deeply examined constraints sit 'bridge' paragraphs stating why progress at one level cannot carry upward, and the non-reducibility claim is what the bridge arguments enforce. A
What would settle it
Give the paper's own test a chance: across ten or more frontier model families evaluated in the same window, if symbol-grounding scores and causal-reasoning scores correlate above 0.7, Prediction 1 is disconfirmed; and a single architectural innovation that independently produces significant gains across three constraints at two or more levels would falsify the general non-reducibility claim outright.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that the constraints any general intelligence must satisfy—symbol grounding at the agent-environment interface, causal reasoning at the computational level, embodiment at the biological level, cultural transmission at the anthropological level, institutional intelligence at the institutional level, and incentive compatibility at the incentive-theoretic level—are mutually non-reducible kinds. Each picks out a distinct property at a distinct level, and no bridge law can state one in the vocabulary of another without losing the generalisations that make it explanatory; the relationship between constraints across levels has the structure the spe
Load-bearing premise
The argument's load-bearing premise is that the multiple-realisability argument from the philosophy of the special sciences extends by analogy to constraints on intelligence that no current system instantiates; if those constraints turn out to be computational requirements described in social vocabulary, the non-reducibility thesis and its anti-scaling conclusion collapse.
Editorial extensions
If this is right
- If the thesis is correct, no continuation of the scaling programme—more data, more compute, larger models—can by itself produce general intelligence.
- No single architectural innovation, whether symbol grounding, causal reasoning, or embodiment, will deliver AGI; progress at one level leaves the other levels untouched.
- AGI research programmes should be evaluated against the full twenty-three-constraint profile, not against performance on any one benchmark family.
- Gains at one level will not transfer to others; the paper points to evidence such as benchmark saturation followed by a reset on a harder successor benchmark as the pattern this predicts.
- Five named falsifiable predictions, each with a disconfirmation condition, convert the taxonomy into a research programme with a longer horizon than the scaling hypothesis.
Reading between the lines
- If the paper is right, current 'AGI' labels based on broad reasoning benchmarks are systematically over-optimistic, because those benchmarks measure at most the computational level of the constraint profile.
- A testable extension suggests itself: track across model families whether gains on one constraint-level benchmark family predict gains on another; low cross-level correlation would support the non-reducibility thesis and high correlation would undermine it.
- The framework implies that AI systems embedded in institutions—participating in legal, market, or governance processes—are not an optional deployment question but part of what general intelligence requires, so an agent trained in isolation would be incomplete by construction.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that general intelligence requires twenty-three structural constraints (eight clusters, six treated in depth) that are mutually non-reducible across distinct levels of description, in the sense of Fodor's (1974) special-sciences argument. The constraints are derived from four evidential lenses — AI systems, anthropology, law, and economics — with speculative fiction used solely as a discovery heuristic. The paper claims that, because the constraints occupy non-reducible levels, no single architectural advance and no continuation of the scaling programme by itself can produce AGI; research programmes should therefore be evaluated against the full constraint profile rather than single benchmarks. The argument issues in five falsifiable predictions with named benchmark families and explicit disconfirmation conditions.
Significance. If the non-reducibility thesis could be sustained, the paper would make a substantive contribution to AGI evaluation, redirecting attention from single benchmarks to a multi-level constraint profile. The paper has real strengths: the demarcation criteria are explicitly stated; the bridge paragraphs in Section 4 are candid about what each constraint does and does not deliver; the predictions in Section 6 are more operational than most framework papers, with named benchmark families and thresholds; the ARC-AGI-1/ARC-AGI-2 and CLadder episodes are accurately described and used in a falsifiable way; and Section 7 acknowledges central limitations, including the sample-of-one problem and the asymmetry in evidential support across constraints. The paper would be a useful framework piece if the gap between explanatory non-reducibility and the anti-scaling conclusion were closed or the conclusion appropriately weakened.
major comments (4)
- [Abstract; §5; §4 bridge paragraphs] The central inference — from Fodorian non-reducibility to 'no single architectural advance, and no continuation of the scaling programme by itself, can produce AGI' — does not follow. Fodor's argument blocks bridge-law reduction and supports explanatory autonomy; it does not entail that one physical/architectural mechanism cannot jointly realize many higher-level kinds, nor that one intervention cannot improve performance at several levels. The bridge paragraphs in Section 4 show only that perfect grounding does not by itself deliver causal reasoning, that causal reasoning does not by itself deliver embodied concepts, and so on. That is insufficiency of one constraint for another, not impossibility of joint implementation. Section 5 explicitly stakes the practical conclusion on explanatory non-reducibility, but explanatory non-reducibility says only that vocabularies and generalisations
- [§2.1; §5] Non-reducibility within levels is partly true by construction. The demarcation criteria (2) 'independence' and (3) 'level specificity' were used in Section 2.1 to merge five candidates and relocate another, reducing the original thirty to twenty-three. Section 5 then states that within-level non-reducibility 'follows from the demarcation criteria of Section 2.1.' This is circular as a proof of non-reducibility: it shows only that the taxonomy contains no redundant entries by its own criteria, not that the constraints are independent Fodorean kinds with proprietary generalisations. The five same-level reduction rebuttals in Section 5 are substantive and helpful, but they do not establish the general claim. In particular, the normative-procedural and incentive-theoretic levels are asserted to be genuinely distinct from the computational level, but the paper concedes that the application to
- [§7; §2.1] The necessity criterion — the instrument that separates what any general intelligence requires from what the human instance exhibits — is acknowledged in Section 7 to be 'easier to state than to discharge.' This is not merely a minor caveat: for several constraints on the ladder, especially embodiment, cultural transmission, and social intelligence, the evidence is drawn from the single human instance, and the paper itself says the necessity claim is less strongly supported for these than for symbol grounding and causal reasoning. Since these are load-bearing rungs of the ladder from which the anti-scaling conclusion is drawn, the conclusion must be correspondingly qualified. The paper partially does this in Section 7, but the abstract and conclusion do not carry the qualification.
- [§6; Table 2] The predictions are presented as making the framework falsifiable, but several disconfirmation thresholds are specified without independent justification and some are difficult to apply as written. For example, Prediction 1 declares disconfirmation at a Spearman correlation exceeding 0.7 across ten model families; Prediction 5 requires a comparison reviewed by two independent groups, which is an editorial procedure rather than a quantitative effect; and Prediction 3's 'proportional gains' criterion uses an arbitrary factor-of-two definition. The paper wisely warns against cross-checkpoint correlations, but without pre-registration or power analysis, the threshold choices look post hoc. This weakens the advertised status of the predictions as a testable research programme, though it does not by itself undermine the taxonomy.
minor comments (5)
- [§2.1; Fig. 1] The number of levels is inconsistent. The text says there are four lenses and four main levels, but Figure 1's caption says 'five levels of description,' and Section 5 adds 'two finer levels' (agent-environment interface, biological/affective). Please reconcile the level taxonomy with the cluster/level entries in Table 1.
- [§4.2] Grammar/agreement: 'CLadder (Jin et al., 2023) show' should be 'shows.' Minor, but the paper is otherwise carefully written.
- [§6; Fig. 6] The phrase 'single digits through about 24 percent' is confusing. The intended meaning is presumably 'from low single digits to roughly 24 percent.' Please rephrase for clarity.
- [Table 2; Prediction 2] The disconfirmation condition for Prediction 2 is stated as causal reasoning correlating 'as strongly ... as the latter correlates with institutional proxies,' but the rationale predicts that cultural-institutional co-variation is higher than either's correlation with causal reasoning. The condition is logically the right null, but the wording should be tightened to make the comparison explicit.
- [Abstract; §5] Key terms 'architectural advance' and 'scaling programme' are never defined. Since the main negative conclusion turns on them, the definitions should be stated early. Also, Figure 5 is explicitly qualitative and scores no actual system, so its use as an 'evaluation template' should be labelled illustrative rather than operational.
Circularity Check
Within-level non-reducibility is enforced by Section 2.1 demarcation criteria; cross-level claim depends on stipulated level assignments.
-
self definitional
[Section 5, third paragraph; Section 2.1 demarcation criteria (2)-(3)]
"Non-reducibility within levels, between constraints at the same level of description, follows from the demarcation criteria of Section 2.1, and the most plausible reduction attempts can be addressed directly."
Section 2.1 retains a candidate only if it satisfies (2) 'independence' and (3) 'level specificity', and merges any candidate that fails them: 'Where a candidate fails criteria (2) or (3), it is merged with the most fundamental surviving constraint.' Every retained constraint is therefore independent and level-specific by stipulation. To say non-reducibility 'follows from' those criteria is to say the taxonomy was constructed so that no two entries are reducible; the conclusion is an entry requirement, not a derived result.
-
self definitional
[Section 2.1, fourth paragraph; Section 5, third paragraph]
"Pairing each lens with a level of description is the load-bearing move of the method, because it is what makes the eventual non-reducibility claim more than an assertion that the disciplines happen to disagree."
The paper's cross-level non-reducibility is obtained by first stipulating that each lens is anchored to a distinct level and that constraints are assigned to those levels, then applying Fodor: 'Because each constraint picks out distinct kinds at a distinct level, the relationship between constraints across levels has exactly the structure Fodor identifies as precluding reduction.' This makes the conclusion follow from the level-assignment, the very point at issue. The bridge paragraphs add substantive considerations, so this is partial rather than total circularity.
full rationale
This paper's central non-reducibility claim is partly circular. Section 5 explicitly says same-level non-reducibility 'follows from the demarcation criteria of Section 2.1,' and those criteria (independence, level specificity) are applied to filter thirty candidates down to twenty-three, merging any candidate that fails them. Retained constraints are therefore non-reducible by construction. The cross-level ladder is also loaded by the method's 'load-bearing move' of pairing each lens with a distinct level of description; the Fodor conclusion is then read off the stipulated level assignment, though the individual bridge paragraphs contain genuine substantive arguments (e.g., grounding does not deliver causation). The five predictions are genuine empirical bets with named benchmarks and disconfirmation thresholds, not fitted parameters, so they are not circular in the statistical sense. There is no self-citation chain or borrowed uniqueness theorem. The anti-scaling conclusion ('no single architectural advance... can produce AGI') overreaches the non-reducibility premise, but that is a logical/correctness issue rather than a circularity. Overall score 6: one central component is definitional, while the empirical/predictive apparatus has independent content.
Assumptions & free parameters
free parameters (3)
- Disconfirmation correlation threshold (Prediction 1) =
Spearman > 0.7 across at least 10 model families
- Proportional-gain factor (Prediction 3) =
within a factor of two
- Significance and sample requirements (Predictions 4 and 5) =
p < 0.05 across at least 5 evaluations; at least 3 constraints, 3 prior-generation models, 2 independent review groups
assumptions (7)
- domain assumption Fodor's (1974) special-sciences non-reducibility: higher-level kinds are multiply realizable, so cross-level bridge laws cannot be natural-kind predicates.
- ad hoc to paper The analogical extension of Fodor's argument from relations among sciences to relations among constraints on intelligence.
- ad hoc to paper The four lenses (AI systems, anthropology, law, economics) are anchored to genuinely distinct Fodorean levels of description.
- ad hoc to paper The necessity demarcation criterion can separate what any general intelligence requires from what the single human instance contingently exhibits.
- domain assumption Working definition of general intelligence as cross-domain skill acquisition (Legg and Hutter 2007 refined by Chollet 2019).
- domain assumption Empirical state-of-the-art claims: frontier LLMs at 'Emerging AGI' level, o3 at about 87.5% on ARC-AGI-1, ARC-AGI-2 best about 24%, CLadder associational-ceiling pattern.
- domain assumption Every retained constraint must be grounded in at least two independent evidential traditions.
invented entities (2)
-
Fodorean levels of description applied to AGI constraints (computational, evolutionary-developmental, normative-procedural, incentive-theoretic)
independent evidence
-
Constraint-profile evaluation template (Fig. 5)
Cite this review
Pith. "Pith review of What General Intelligence Requires: Non-Reducible Constraints Across Levels of Description." pith.science (2026). https://pith.science/paper/QZA7ECCH
@misc{pith2026260718943,
author = {Pith},
title = {Pith review of: What General Intelligence Requires: Non-Reducible Constraints Across Levels of Description},
year = {2026},
howpublished = {\url{https://pith.science/paper/QZA7ECCH}},
note = {Machine review of arXiv:2607.18943}
}
read the original abstract
General intelligence, of the kind that underwrites the full range of human cognitive achievement, is not a property of computational architecture alone. This paper advances a single thesis: the structural constraints on general intelligence occupy distinct levels of description and are mutually non-reducible, in the sense that the special-sciences tradition gives to that term. It follows that no single architectural advance, and no continuation of the scaling programme by itself, can produce artificial general intelligence (AGI), and that research programmes must be evaluated against the full constraint profile rather than against performance on any one benchmark. The thesis is developed through a method that reads general intelligence through four evidential lenses, AI systems research, anthropology, law, and economics, each anchored to a distinct level of description, supplemented by speculative fiction used as a disciplined heuristic in the context of discovery rather than the context of justification. Applying the method yields a taxonomy of twenty-three structural constraints organised into eight clusters; six are examined in depth and ordered as an ascending ladder of levels, with explicit bridges showing why progress at one level cannot carry to the next. The argument issues in five falsifiable predictions, each stated with a named benchmark family and a disconfirmation condition, converting a descriptive framework into a research programme with a longer horizon than the scaling hypothesis implies.
Reference graph
Works this paper leans on
-
[1]
Acemoglu, D. (2023). Distorted innovation: Does the market get the direction of technology right? AEA Papers and Proceedings, 113, 1-28. https://doi.org/10.1257/pandp.20231000 Acemoglu, D., & Johnson, S. (2023). Power and progress: Our thousand-year struggle over technology and prosperity. PublicAffairs. Acemoglu, D., & Restrepo, P. (2020a). The wrong kin...
arXiv 2023
-
[4]
https://arxiv.org/abs/2303.12712 Chollet, F
arXiv:2303.12712. https://arxiv.org/abs/2303.12712 Chollet, F. (2019). On the measure of intelligence. arXiv:1911.01547. https://arxiv.org/abs/1911.01547 70 Chollet, F., Knoop, M., Kamradt, G., Landers, B., & Pinkard, H. (2025). ARC-AGI-2: A new challenge for frontier AI reasoning systems. arXiv:2505.11831. https://arxiv.org/abs/2505.11831 Clark, A., & Ch...
arXiv 2019
-
[30]
H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., & Fedus, W
https://arxiv.org/abs/1706.03762 Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., & Fedus, W. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research. https://openreview.net/forum?id=yzkSU5...
arXiv 2022
-
[36]
https://arxiv.org/abs/2312.04350 Kahneman, D. (2011). Thinking, fast and slow. Farrar, Straus and Giroux. Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling laws for neural language models. arXiv:2001.08361. https://arxiv.org/abs/2001.08361 Kim, M. J., Pertsch, K., Kara...
arXiv 2011
-
[2025]
https://arxiv.org/abs/2410.24164 69 Borges, J. L. (1941/1944). The library of Babel. In Ficciones. Editorial Sur. (Original story first published 1941; republished in Ficciones, 1944.) Brooks, R. A. (1991). Intelligence without representation. Artificial Intelligence, 47(1-3), 139-159. https://doi.org/10.1016/0004-3702(91)90053-M Brown, T. B., Mann, B., R...
arXiv 1941
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.