{"id":"2eda1157-e8d6-4497-b8cf-076e8effc221","arxiv_id":"2509.11131","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review arguing that trainable cellular automata form a unifying, computationally lean paradigm for biology-inspired collective intelligence and generative AI.","lead":"This paper reviews neural cellular automata, a type of AI where each cell of a grid has a small neural network and the whole grid learns to grow, repair, or control patterns and robots. It argues that these systems form a unifying, computationally lean framework connecting biology, collective intelligence, and modern generative AI.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ARC-AGI-1 'competitive with ChatGPT-4.5' claim is the only quantitative evidence for NCAs beyond classical AI, and it is unverified in the review; an independent audit of refs 24/25 should settle it.","rationale":"The reader's UNVERDICTED assessment is appropriate, and my concern does not change it: this is a review paper whose central claims depend on cited primary results that are not reproduced or audited here. Among those cited results, the ARC-AGI-1 comparison is the most load-bearing because it is the only quantitative, task-level evidence for the 'beyond classical AI' part of the central claim. The aging/regeneration simulations are also cited, but they support the biology side, which the review itself frames as toy modeling; a single overstated simulation there would be less harmful to the central paradigm claim than an overstated reasoning benchmark. I am not alleging misconduct; the issue is that a review paper's central claim should not rest on unverified headline numbers. The §5 limitations are honest, but they are in tension with the Abstract's confidence: if NCAs are toy-models with storage and scale limits, 'competitive with ChatGPT-4.5' needs clear technical qualification. My proposed audit would settle whether the flagship quantitative claim actually holds.","tokens_in":17548,"tokens_out":5452,"duration_ms":64021,"concrete_test":"Run an independent audit of refs 24 and 25 against the official ARC-AGI-1 evaluation protocol: obtain the released code/checkpoints, reproduce the solve rate on the full public test split (or the exact split used in each paper), and locate the ARC Prize leaderboard entry for ChatGPT-4.5 using the same number of examples per task. Compare solve rates and total compute, including NCA training, evolution/search, and inference across cells. If the reproduced NCA solve rate is within sampling error of ChatGPT-4.5's and cost is lower after full accounting, the §4 claim stands; if the reported comparison used a nonstandard split, cherry-picked tasks, or omitted the search/training compute, the paper should weaken the Abstract to 'promising preliminary results' rather than 'competitive.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing issue is the ARC-AGI-1 comparison in §4 and the Abstract. The paper asserts that 'vanilla NCAs and EngramNCAs perform competitively on the ARC challenge compared to ChatGPT-4.5' and uses this as flagship evidence that NCAs go 'beyond classical AI' toward hierarchical reasoning and control. No solve rates, task subset, evaluation protocol, cost figures, or confidence intervals are reported; the claim rests entirely on two arXiv preprints (refs 24, 25) that are not independently verified in the manuscript. This is an external-evidence dependence, and it is load-bearing because the Abstract explicitly names ARC-AGI-1 as the reasoning achievement of the paradigm. The paper's own §5 cautions that NCAs are 'toy-models' and that scaling and storage capacity are unsolved, so the strong ARC claim is not supported by the internal evidence. If the ARC results are materially overstated or not directly comparable to ChatGPT-4.5 (e.g., different number of demonstrations, selection of easy tasks, or evolutionary search counted outside the reported cost), the 'unifying computationally lean paradigm' claim loses its only quantitative demonstration of reasoning. This is a correctness risk in the review's central claim, not a disagreement with consensus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a review/position paper on Neural Cellular Automata (NCA) in biology and AI. It surveys recent work on NCAs for morphogenesis, regeneration, aging, bioelectricity, evolution, the generative genome, molecular design, and robotics, and argues that NCAs instantiate a 'multiscale competency architecture' analogous to biological organization. It further claims that NCAs go beyond classical AI: they show robustness, generalization, decentralized control, and even competitive performance on ARC-AGI-1 abstraction/reasoning tasks (compared with ChatGPT-4.5) at a fraction of cost. The paper connects NCA dynamics to diffusion models and criticality, and concludes by advocating NCAs as a unifying 'computationally lean' paradigm for bio-inspired collective intelligence. The review is broad and cites extensive literature, but its flagship quantitative claims about ARC are not substantiated within the manuscript and rest on arXiv preprints [24,25]; the paper's own limitations section acknowledges NCAs are toy models with unsolved scaling and storage issues.","tokens_in":17881,"tokens_out":6199,"duration_ms":75215,"significance":"If the reported ARC-AGI-1 results and the generalization-from-few-examples claims were verified, the paper would point to a substantial result: NCAs as a low-cost, decentralized alternative to large transformer models on abstract reasoning. The review also usefully consolidates the diverse NCA literature and gives credit to a self-contained list of applications and open problems, including criticality and hybrid evolution/diffusion training. However, as a review it contributes no new experiments or formal analysis; its significance therefore rests on the accuracy of cited prior work and on the coherence of the 'multiscale competency' framing. The authors are explicit about many limitations, which is commendable and should be retained.","major_comments":[{"comment":"The sentence 'vanilla NCAs and EngramNCAs perform competitively on the ARC challenge compared to ChatGPT-4.5 – notably at the fraction of costs and computational resources' is a quantitative, load-bearing claim, but no numbers or protocols are given. The manuscript does not report solve rates, task subset, number of demonstrations, inference budget, or what 'competitive' means in this context; it only cites refs [24,25], both arXiv preprints. Likewise, 'can learn to generalize ... from a minimal set of two or three training examples' is stated without supporting evaluation. Since the Abstract names ARC-AGI-1 as the main evidence for NCAs 'beyond classical AI', this unverified external dependence must be resolved. At minimum, state the reported accuracies and evaluation conditions, or explicitly label the claim as reported by the cited preprints with appropriate caveats.","section":"§4 (ARC-NCA paragraph) and Abstract"},{"comment":"The limitations section concedes NCAs 'represent toy-models for biological organization', that simulating realistic organ-level complexity at unicellular resolutions is 'currently infeasible', and that integrating molecular/genetic/biomechanical detail is difficult. The conclusion nevertheless calls NCAs 'a model of choice' for AI-oriented computational biology and the basis of a 'unifying computationally lean paradigm'. A review may advocate in spite of limitations, but the manuscript should specify the intended scope (e.g., conceptual models of collective dynamics rather than quantitative organ simulation) and explain why the identified gaps do not prevent the stated 'model of choice' status. As written, the conclusion outruns the evidence assembled in the paper.","section":"§5 vs. §6"},{"comment":"The 'multiscale competency architecture' is the central interpretive frame, but it is never defined operationally. The reader cannot tell whether a given NCA 'has' a competency at a scale, how competencies are measured, or how hierarchical NCA layers (refs [10,11]) implement 'competency amplification' as opposed to a generic multi-scale model. Since the unifying claim depends on this equivalence, the authors should provide a concrete criterion or a worked example, mapping one NCA architecture cell-by-cell to homeostatic loops at each level. Otherwise the central thesis remains resistant to evaluation.","section":"§1, §4: multiscale competency architecture"}],"minor_comments":[{"comment":"The mathematical notation is garbled in the provided text (e.g., 's!\"#$=f/s!\",{s%\"}%∈𝒩(𝒾)2' and 'x!\"#$=x!\"+Δx!\"\"'). Please check the typesetting and define all symbols consistently.","section":"Section 2, equations"},{"comment":"Typos and OCR artifacts: 'is constraint to' should be 'is constrained to'; 'William Jame's' should be 'William James's'; reference list contains 'diWerentiable', 'parameter-eWicient', 'DiWusion' instead of 'Differentiable', 'parameter-efficient', 'Diffusion'.","section":"Abstract and throughout"},{"comment":"The sentence 'the NCA's grid stats serve as recurrent feedback signal' should read 'grid states'. Also, the claim that dynamics are 'differentiable across temporal state updates' should be reconciled with the asynchronous/stochastic updates described later in the same section.","section":"Section 2"},{"comment":"Some panels (e.g., K, referencing ARC-AGI-1 tasks) are cited without a quantitative caption. Consider adding one line describing what is shown and the reported performance.","section":"Figure 1"},{"comment":"The paragraph on modularity/interfacing is hard to parse: 'this limits modularity and compatibility across multiple NCAs that operate in the same environment but developed not necessarily compatible communication strategies'. Please rephrase for clarity.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The review leans heavily on the authors' own prior work and on two arXiv preprints (refs 24,25) for the strongest claim about ARC-AGI-1. I recommend that the editor request either an independent audit of the ARC numbers or an explicit softening of the claim to 'as reported in refs [24,25]' before publication. The paper may be better framed as a perspective/review than as a primary research contribution; the ARC claim should not be published without the underlying quantitative details."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a genuinely useful review of neural cellular automata in biology, and you can trust its literature coverage, but its central 'beyond classical AI' claim—competitive ARC-AGI-1 performance—is carried entirely by two preprints and not checked in the article. Read those preprints before you repeat the claim.\n\nWhat the paper does well: it organizes a large and scattered literature into a coherent picture, connecting NCA mechanisms to Levin's multiscale competency framework without getting lost in jargon. The sections on regeneration, aging, and the diffusion analogy are clear and appropriately hedged. The limitations section deserves credit: it explicitly says NCAs are toy-models, that scaling to organ-level biology is infeasible, and that storage of multiple attractors is unsolved. That honesty is rare in a field that tends to oversell.\n\nThe soft spots are proportionate to the review format. No new results, which is fine for a review. The bigger issue is that the abstract and Section 4 sell NCAs as a unifying paradigm based on the ARC-AGI-1 comparison, but give no solve rates, no protocol, no cost figures. It rests on refs 24 and 25, current preprints from overlapping groups. The stress test is right to call this load-bearing: if those preprints don't hold up, the paper's main novelty—NCAs on reasoning tasks—collapses to 'interesting but unverified.' The authors could fix this with a small table of solve rates and an explicit caveat about comparability (e.g., number of demonstrations, whether evolution is counted).\n\nOne more note: the citation pattern leans heavily on the authors' own papers. Not damning—they did much of this work—but the narrative would be stronger with more independent validation.\n\nBottom line: it deserves peer review, but a referee should push on the ARC numbers before acceptance. I'd use it as a review reference myself, and I'd bring it to a reading group interested in the biology-AI interface.","headline":"A useful survey of NCAs in biology, but the flagship ARC-AGI-1 claim is carried on preprints the paper never verifies.","tokens_in":18275,"tokens_out":2903,"would_cite":true,"duration_ms":34075,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68Q80","92C15","68T05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that neural cellular automata constitute a unifying, computationally lean paradigm spanning biological self-organization, distributed robotic control, and abstract reasoning, potentially enabling a new class of bioinspire","keywords":["neural cellular automata","collective intelligence","multiscale competency","morphogenesis","regeneration","ARC-AGI","diffusion models","distributed control"],"falsifier":"Independently reproduce the ARC-NCA and EngramNCA runs on a held-out subset of ARC-AGI-1 under the stated training budget and compare with a large language model baseline; if the performance gap is much smaller or the cost estimate is off by orders of magnitude, the strongest claim fails. Alternatively, show that NCA-trained morphologies cannot regenerate after damage when scaled beyond toy grid sizes, contradicting the claimed scalability.","tokens_in":17464,"feed_emoji":"🧬","tokens_out":4584,"duration_ms":51684,"temperature":0.7,"pith_summary":"This review argues that neural cellular automata—grids of locally interacting cells, each governed by a small trainable neural network—form a single computational paradigm that spans biological self-organization, distributed robot control, and abstract reasoning. The authors contend that this architecture mirrors the multiscale competency organization of living systems, where nested levels of agents pursue local goals that collectively produce system-level outcomes. They point to results showing that NCAs regenerate damaged patterns, model aging and bioelectricity, control soft robots, and perform competitively on the ARC-AGI-1 reasoning benchmark at a small fraction of typical costs. If the argument holds, NCAs would offer a biologically grounded alternative to large, centralized deep-learning models.","feed_headline":"Local cell rules could span biology, robots, and reasoning","feed_subtitle":"A review argues these trainable cellular automata form a lean bridge from morphogenesis to ARC-AGI.","key_machinery":"The central object is the neural cellular automaton (NCA): each cell on a discrete grid keeps a continuous vector state and updates it using a shared, small neural network that sees only the local neighborhood—normally a 3×3 convolution plus a dense layer, on the order of 10,000 parameters. Training is done by gradient descent or evolution on global losses, e.g., matching a target image or solving a maze, while stochastic updates and damage-masking during training encourage regeneration. The paper also highlights architectural extensions: private engram channels for memory and transfer, hierarchical stacked NCAs for multi-scale coupling, and criticality-pretrained cells. These make the NCA a","core_discovery":"The core claim is that the NCA architecture—a spatial grid in which every cell runs the same feed-forward neural network update rule on local neighbor states—is not just a tool for pattern formation but a substrate for collective intelligence. The authors survey evidence that these systems self-assemble and repair morphologies, store and transfer genetic information through private cell-state channels, regenerate 3D machines, and even grow solutions to ARC-AGI-1 tasks from a few examples. They argue that the same iterative, locally constrained refinement that makes NCAs robust also links them to denoising diffusion models, and that the architecture's built-in multiscale competency makes it a","pith_inferences":["An independent replication of the ARC-NCA and EngramNCA runs on a held-out ARC-AGI-1 subset would sharply test the strongest claim; the review does not audit those results itself.","If the diffusion analogy is more than superficial, NCA training could benefit from explicit time-conditioning or noise-schedule curricula without giving up locality—a testable design change.","The multiscale-competency perspective predicts that NCA-like systems will show better damage tolerance and transfer when trained with hierarchical or private-state architectures; one could measure this directly against vanilla NCAs.","The toy-model caveat suggests the field needs benchmarks at organ-level complexity before the unification claim is secure."],"forward_implications":["If NCAs remain competitive on ARC-AGI-1 at a fraction of the cost, abstraction and reasoning tasks may be approachable with much smaller, decentralized models.","NCA-based models of regeneration and aging suggest that tissue-level goal-directedness, not just cellular damage, is a key variable in decline—implying targeted interventions can reactivate dormant regenerative potential.","Since NCAs double as distributed controllers for soft and voxel-based robots, the same learned morphogenetic rules could give machines self-repair and adaptation capabilities.","The structural parallel between NCA refinement and diffusion denoising suggests a path to hybrid generative models that keep spatial locality and parameter efficiency.","A biology-inspired AI built on NCAs would be modular and interpretable with the tools of neuroscience and psychology, potentially easing alignment problems."],"fun_headline_variants":["Decentralized cell rules regenerate robots and reason","Same local rule builds bodies, robots, and AI","Trainable cell grids span morphogenesis to ARC-AGI","From morphogenesis to robots: one neural rule","Single update rule powers biology, robots, and reasoning"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole synthesis depends on the accuracy and representativeness of the primary results it cites—especially the reported ARC-AGI-1 performance of NCAs and the aging/regeneration simulations—since the paper reviews rather than reproduces them.","fun_headline_variants_meta":{"raw":{"variants":["Decentralized cell rules regenerate robots and reason","Same local rule builds bodies, robots, and AI","Trainable cell grids span morphogenesis to ARC-AGI","From morphogenesis to robots: one neural rule","Single update rule powers biology, robots, and reasoning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000553,"raw_usage":{"total_tokens":2502,"prompt_tokens":806,"completion_tokens":1696,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":1620}},"tokens_in":550,"tokens_out":1696,"duration_ms":15381,"temperature":1.0,"reasoning_tokens":1620,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T17:02:09.629165+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently reproduce the ARC-NCA and EngramNCA runs on a held-out subset of ARC-AGI-1 under the stated training budget and compare with a large language model baseline; if the performance gap is much smaller or the cost estimate is off by orders of magnitude, the strongest claim fails. Alternatively, show that NCA-trained morphologies cannot regenerate after damage when scaled beyond toy grid sizes, contradicting the claimed scalability.","supporting_citations":[],"review_version":1}