{"id":"77de79e9-5f31-4bd7-bcac-e5641207c0f4","arxiv_id":"2504.20197","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A percolation model of data distributions predicts context, component, and surface features, with sparse or dense regimes depending on occupation probability.","lead":"This paper models a neural network's data distribution as a random lattice and uses percolation theory to explain why learned features come in three types. The model gives qualitative predictions for when features are sparse or dense, and why autoencoder features often split or form hierarchies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3.4's inference from percolation fractal dimension D=4 to four-dimensional component-feature representations is an unsupported leap; no mechanism links input-space cluster geometry to activation-space dimensionality.","rationale":"The reader's weakest assumption identifies the same gap: Section 3.4 moves from D=4 (a fractal dimension of clusters in the input lattice) to four-dimensional neural representations without a mechanism. My stress-test agrees this is the most load-bearing point. The appendix derivation of D=4 is internally sound as a percolation-theory calculation; the issue is the transfer to representation dimension. The broader taxonomy (context/component/surface features, feature sparsity by regime, hierarchical feature splitting) is qualitative and could survive even if the 4D prediction fails; conversely, the 4D prediction is the paper's only quantitative, falsifiable claim. Because the reader already graded this as high-correctness-risk and conditional, my assessment does not move the verdict: it remains CONDITIONAL (recorded as UNCHANGED). The paper should either derive a concrete information-theoretic or geometric mechanism linking cluster geometry to feature-space dimensionality, or soften the 4D claim to a heuristic analogy.","tokens_in":11247,"tokens_out":10233,"duration_ms":120543,"concrete_test":"Analytic falsification test: apply the Section 3.4 inference to a one-dimensional random-walk path in R^6. Such a path has Euclidean fractal dimension 2, so the paper's rule would predict two-dimensional component features; but the path is faithfully represented by a single arc-length coordinate. If the paper's inference is valid, it must specify why percolation clusters differ; this minimal case shows that input-space fractal dimension alone does not fix representation dimensionality. To make it quantitative, compute the minimal k for which the component function on a critical percolation cluster (Bethe lattice, z=6, at p_c) can be written as g(c_1(x),...,c_k(x)) with continuous features c_i; if k<4 for large clusters, the D=4-to-dimension claim is not a mathematical consequence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is in Section 3.4: 'component features should generically require four-dimensional representations, as they map out treelike data volumes that have four-dimensional fractal geometry.' Appendix A derives D=4 as the Euclidean fractal dimension of critical percolation clusters for d>=6 (mass within Euclidean radius r scales as r^4). But D is a property of the cluster embedded in the input lattice. A neural representation is a learned coding scheme and need not preserve input-space Euclidean distances or volume-radius scaling; nothing in Sections 2-3 or Appendix A shows that a 4-dimensional activation subspace is necessary or sufficient to compute a cluster's target function. The cluster topology is treelike and can be navigated by a depth coordinate plus branch labels. A one-dimensional random-walk path embedded in R^6 has Euclidean fractal dimension 2 but can be represented by one arc-length coordinate, showing that fractal dimension of an embedding does not constrain coding dimension. The paper's most concrete quantitative prediction therefore rests on an analogy, not a derivation. If this link fails, the 'four-dimensional representations' claim collapses, although the context/component/surface taxonomy and the qualitative percolation picture could survive.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes modeling a generic data distribution as a site-percolation model on a d-dimensional lattice. It reviews percolation theory (Appendix A) to obtain cluster size distribution exponent tau=5/2 and fractal dimension D=4 for d>=6. It classifies learned features into context, component, and surface features, and argues that this taxonomy and the percolation geometry explain observations from mechanistic interpretability, including feature sparsity, hierarchy, feature splitting, and multidimensional features. It predicts that component features generically require four-dimensional representations, and that feature usage has a power-law distribution.","tokens_in":11507,"tokens_out":8515,"duration_ms":83079,"significance":"If the proposed link between percolation geometry and learned representations were established, this would provide a unifying geometrical foundation for feature decompositions in neural networks, connecting data distributions to representation structure. The paper brings a standard statistical-physics tool into mechanistic interpretability and makes one concrete, falsifiable prediction (four-dimensional component features). Its strengths include a correct and clearly presented review of high-dimensional percolation exponents in Appendix A, and a taxonomy that plausibly organizes several existing interpretability findings. However, the central inference from percolation clusters to neural features is asserted rather than derived, and the most quantitative prediction rests on an analogy between input-space fractal dimension and activation-space dimensionality. The paper therefore reads as a promising framework or position paper rather than an established theory.","major_comments":[{"comment":"The claim that component features \"should generically require four-dimensional representations, as they map out treelike data volumes that have four-dimensional fractal geometry\" does not follow from the derived D=4. D is the Euclidean fractal dimension of a percolation cluster embedded in the input lattice; the dimensionality of a neural representation is a property of the network's coding scheme, and no equation or mechanism connects the two. A concrete counterexample: a one-dimensional random-walk path embedded in R^6 has Euclidean fractal dimension 2 but can be encoded with a single arc-length coordinate, showing that fractal dimension of an embedding does not constrain coding dimension. Unless a mechanism is supplied (e.g., a proof that computing a cluster's target function requires a representation whose dimension equals the cluster's fractal dimension), this most concrete quantitative prediction is an analogy, not a consequence of the model.","section":"Section 3.4, with Appendix A (Eqs. 12, 18)"},{"comment":"The mapping from percolation clusters to features is not formalized. The model defines clusters via a bond rule on the target function (Section 2.1) and then asserts that each cluster supports a distinct target function and that context/component/surface features are \"a minimal set of sparse or composable latent variables\" (Section 3.1). No equations connect cluster statistics (n_s, D) to feature count, feature dimensionality, or feature activation frequency. For example, the statement in Section 3.4 that \"both context and component features should have a power-law distribution in use frequency\" is not derived from n_s ~ s^{-tau}; it is not even specified whether feature frequency is proportional to cluster size, cluster volume, or some other observable. The central chain from data geometry to feature statistics is therefore underived.","section":"Sections 2.1-3.3"},{"comment":"The three-category taxonomy appears to be introduced to accommodate existing interpretability findings (Gurnee et al., Bricken et al., Engels et al., Ilyas et al.) rather than derived from the percolation model. Because the model has unconstrained parameters (p, d) and no independent criterion for which regime applies, its qualitative predictions can accommodate both sparse and dense features, and both many and few features, by tuning p. For instance, the classification of modular addition as a \"surface feature\" in Section 3.1 is purely verbal and not tied to any property of the percolation model. Without a calibration of p to measurable dataset characteristics (which the paper itself notes is needed in the footnote to Table 1), these are not falsifiable predictions.","section":"Section 3.1 and Table 1"},{"comment":"The model assumes both that the data distribution is near the percolation threshold (p ≈ pc) and that the occupied-site process is independent site percolation. These are substantive assumptions, not consequences of the target-function bond rule. All the paper's quantitative claims (power-law cluster size distribution, D=4) hold only at criticality in the scaling limit, yet no argument is given for why a generic data distribution should be critical. The paper should either derive criticality from a plausible mechanism (e.g., a maximum-entropy or self-organized criticality argument) or explicitly present it as a falsifiable assumption with observable consequences, rather than treating it as the default case.","section":"Section 2.1"}],"minor_comments":[{"comment":"The sentence contrasting the predicted four-dimensional component features with the \"circular feature manifolds\" of Engels et al. (2024a) is confusing: the paper claims to be \"in line with\" Engels et al. while simultaneously predicting a different geometry. Clarify how the two claims relate.","section":"Section 3.4"},{"comment":"The sentence \"This sparsity suggests that models may efficiently represent these features polysemantically, in superposition Elhage et al. (2022)\" is missing a comma before the citation, making it read as if superposition is the name of a paper rather than a concept.","section":"Section 3.3"},{"comment":"The notation sigma=1/2 is introduced in Eq. 8 and used in later exponents; defining sigma as a standard percolation exponent at its first mention would help readers not familiar with the notation.","section":"Appendix A, Eq. 8"},{"comment":"Several references are to blog posts and LessWrong posts (Nabeshima 2024, Bussmann et al. 2024, Mendel 2024, Yudkowsky 2008). These are appropriate for a fast-moving field, but the journal's reference guidelines should be checked for consistency.","section":"References"},{"comment":"The paper criticizes Elhage et al.'s feature definition as circular, but its own definition of a feature (\"coordinates for computing the target function\") is not stated formally. A precise definition of 'feature' would strengthen the claims.","section":"Section 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a theory/position paper with no experiments. Its appendix is a correct review of standard high-dimensional percolation results, but the novel part—the mapping from percolation clusters to learned features—is asserted rather than derived. The concrete prediction of four-dimensional component features is the most likely point to be cited, yet it is currently an unsupported analogy. I recommend major revision: the authors should either derive the claimed link, or substantially soften the claim and reframe the paper as a testable framework with explicit open problems. There is also a notable pattern of self-reference (the model is borrowed from the author's Brill 2024), which is not inherently problematic but should be acknowledged and, if possible, validated with external evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things about this paper. First, it is a genuinely readable, candid proposal that uses percolation theory to give a geometric account of feature taxonomy and SAE phenomena. Second, its headline quantitative prediction, that component features generically require four-dimensional representations (Section 3.4), is not derived from the percolation math. It is an analogy, and the authors do not supply a mechanism linking input-space fractal dimension to activation-space dimensionality.\n\nWhat is actually new: the context/component/surface feature taxonomy, and the use of percolation cluster statistics to explain feature splitting, hierarchy, and conditional activation in sparse autoencoders. The appendix derivation of tau=5/2 and D=4 is standard and correct, and the writing is refreshingly honest—the abstract says \"qualitatively consistent,\" which it is. The paper also does a good job of situating itself against the existing SAE literature and acknowledging where formalization is missing (e.g., the meaning of p is left loose).\n\nThe soft spots are real but localized. The D=4 to four-dimensional-representation step is the load-bearing bridge, and it fails under inspection. Fractal dimension of an embedded cluster is a property of the embedding; a neural representation is a coding scheme and need not preserve volume-radius scaling. The stress-test example is apt: a one-dimensional random walk embedded in R^6 has Euclidean fractal dimension 2 but can be encoded with one arc-length coordinate. So the four-dimensional claim is not merely unproven—it is likely false as stated, unless a very specific and unmotivated constraint on the representation is added. The broader taxonomy and the qualitative percolation picture could survive without it, but the paper as written leans heavily on that specific number.\n\nAlso note that the random lattice model itself is taken from the author's own 2024 paper, so the novelty is in the application and taxonomy, not the underlying model. That is fine, but it means the paper is a conceptual reframing rather than a new empirical or mathematical result. No data are shipped, no code, no falsifiable predictions other than the D=4 claim, which is currently a free parameter in disguise.\n\nWho should read this: people working on mechanistic interpretability, especially SAE evaluation and feature geometry. It deserves a serious referee, but the referee should press hard on Section 3.4. My recommendation is to accept for peer review, with the expectation that the authors either derive a credible link or explicitly downgrade the four-dimensional claim to a heuristic.\n\nBest,","headline":"A clear and honest conceptual paper that borrows percolation theory from the author's prior work, but its most concrete claim—that component features need four-dimensional representations—is an analogy, not a derivation.","tokens_in":761,"tokens_out":2030,"would_cite":false,"duration_ms":33869,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["82B43","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that the geometry of a generic data distribution—modeled as a random lattice—forces component features in neural networks to be four-dimensional, because percolation clusters have fractal dimension 4.","keywords":["representation learning","percolation theory","fractal dimension","mechanistic interpretability","sparse autoencoders","feature geometry","data distribution","latent features"],"falsifier":"Find a well-isolated component feature in a trained transformer—one that tracks a compositional relation within a cluster—and measure the intrinsic dimension of its activation subspace; the paper predicts $4$, so a stable intrinsic dimension clearly different from $4$ (for example, the 2-dimensional circular features already reported for days of the week) would refute the central prediction.","tokens_in":11035,"feed_emoji":"🧩","tokens_out":11310,"duration_ms":98065,"temperature":0.7,"pith_summary":"The paper proposes that a generic data distribution, as seen by a general-purpose learner, is well modeled by a random lattice of occupied and empty sites, and that the geometry of the resulting percolation clusters dictates what internal features a neural network must learn. On this model, learned features fall into three categories: context features that identify clusters, component features that coordinate within a cluster, and surface features that arise when the learner exploits format or memorized structure. Its most concrete prediction is that component features should generically occupy four-dimensional representation subspaces, because the treelike percolation clusters they map out have fractal dimension $D=4$ in high-dimensional lattices. A sympathetic reader would care because the model offers an architecture-independent origin story for feature structure and explains why sparse autoencoders see hierarchical, splittable features.","feed_headline":"Percolation theory predicts four-dimensional neural features","feed_subtitle":"A random-lattice model of data distributions derives feature types and dimensionality from percolation cluster geometry.","key_machinery":"The engine of the argument is site percolation on a $d$-dimensional hypercubic lattice, analyzed through the exactly solvable Bethe-lattice approximation for $d\\ge 6$. The load-bearing outputs are the cluster size distribution $n_s \\propto s^{-5/2}e^{-cs}$ and the fractal dimension $D=4$, derived by equating the correlation-length exponents in chemical and Euclidean distance using the random-walk relation $r^2\\propto l$. These cluster statistics are then read as feature statistics: cluster identity becomes context features, within-cluster coordinates become component features, and departures from ideal general-purpose learning become surface features.","core_discovery":"The paper's central claim is that a generic data distribution, viewed through the lens of a general-purpose learner, has the geometry of a random percolation cluster, and that this geometry fixes the structure of the features a network must learn. At or near the percolation threshold in high dimension ($d\\ge 6$), finite clusters are treelike, have a power-law size distribution with exponent $\\tau=5/2$, and have fractal dimension $D=4$. From this the paper concludes that context features identify clusters and their fractal substructure, component features provide coordinates inside a cluster, and component features should generically require four-dimensional representations. This is presented as a quantitative, architecture-independent alternative to the linear representation hypothesis, qualified as consistent with, rather than proven by, existing mechanistic interpretability findings.","pith_inferences":["The author leaves implicit that the $D=4$ prediction is testable without knowing the full data distribution: one can measure the intrinsic dimensionality of component-feature subspaces directly in trained models, and the predicted universality across architectures and modalities would be a strong signal.","A natural extension is to replace hard bonds with weighted bonds, turning crisp percolation clusters into graded family-resemblance concepts; this suggests testable predictions about how feature hierarchies soften as cluster membership becomes graded.","If the random-input-format assumption fails for modalities with strong built-in symmetries, the 4D prediction might hold for high-level semantic features while low-level features track the modality's own geometry, implying that tests should stratify by layer or feature type."],"forward_implications":["Sparse dictionary learning should be most effective for datasets in the subcritical regime; near and above the percolation threshold, features should increasingly activate in dense composition rather than sparse superposition.","Component features should be multidimensional rather than one-dimensional, so interpretability tools that search for 1D linear features will miss part of the structure.","Feature splitting and absorption seen in larger sparse autoencoders reflect an underlying nested hierarchy of context features, not merely a failure of dictionary learning.","The relative numbers of context and component features shift with the occupation regime: many of both below threshold, few of both just above, and only component features when the infinite cluster becomes Euclidean."],"supporting_citations":[{"why":"This work supplies the random-lattice model of data distributions whose geometry the paper analyzes.","marker":"Brill (2024)"},{"why":"This is the source of percolation cluster statistics, critical exponents, and the Bethe-lattice treatment.","marker":"Stauffer and Aharony (1994)"},{"why":"This provides the background for fractal cluster geometry used in the derivation of $D=4$.","marker":"Bunde and Havlin (2012)"},{"why":"This supplies empirical evidence for multidimensional features in language models and the circular feature manifolds contrasted with the predicted 4D geometry.","marker":"Engels et al. (2024a)"},{"why":"This defines the superposition hypothesis and feature definitions that the sparse-autoencoder section evaluates.","marker":"Elhage et al. (2022)"},{"why":"This reports feature splitting in larger sparse autoencoders, cited as evidence for hierarchical context features.","marker":"Bricken et al. (2023)"},{"why":"This identifies monosemantic high-level context features in LLMs, supporting the context-feature category.","marker":"Gurnee et al. (2023)"},{"why":"This contributes the notion of each cluster supporting a distinct target-function 'quantum'.","marker":"Michaud et al. (2024)"}],"fun_headline_variants":["Percolation theory predicts four-dimensional features","Random lattice geometry yields 4D neural representations","Neural features get dimension from percolation clusters","Four-dimensional features from a random lattice model","Percolation explains feature dimensionality in networks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the fractal dimension of percolation clusters in input space ($D=4$) directly forces the number of coordinates a network must use to represent a component feature to be four; the paper does not show a mechanism by which data-space cluster geometry constrains the coding dimension of activations.","fun_headline_variants_meta":{"raw":{"variants":["Percolation theory predicts four-dimensional features","Random lattice geometry yields 4D neural representations","Neural features get dimension from percolation clusters","Four-dimensional features from a random lattice model","Percolation explains feature dimensionality in networks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000209,"raw_usage":{"total_tokens":1317,"prompt_tokens":767,"completion_tokens":550,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":383,"completion_tokens_details":{"reasoning_tokens":481}},"tokens_in":383,"tokens_out":550,"duration_ms":5348,"temperature":1.0,"reasoning_tokens":481,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:34:45.619386+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find a well-isolated component feature in a trained transformer—one that tracks a compositional relation within a cluster—and measure the intrinsic dimension of its activation subspace; the paper predicts $4$, so a stable intrinsic dimension clearly different from $4$ (for example, the 2-dimensional circular features already reported for days of the week) would refute the central prediction.","supporting_citations":[],"review_version":1}