{"id":"a9732a33-17b2-4b75-92e9-3f9bdff3a116","arxiv_id":"2607.10034","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A closed-form Transformer-compatible MLP stores facts at optimal information-theoretic capacity, using 10–104× fewer parameters than prior constructions under isotropic embeddings.","lead":"The authors give a closed-form construction of MLP layers that store facts at information-theoretically optimal capacity and plug into Transformers. If correct, this explains why LLMs pack facts so densely and enables smaller fact modules and modular fact editing.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Optimality and Transformer compatibility both rest on decoding-margin conditions that the abstract asserts but never defines or states precisely enough to check.","rationale":"The reader’s weakest_assumption correctly isolates the decoding-margin analysis as the linchpin that is left unspecified. With only the abstract in hand, no deeper technical soft spot can be diagnosed; the concern is exactly the one already flagged. No proofs, code, tables or formal statements are present that could raise or lower the risk assessment, so the UNVERDICTED / LOW-confidence stance remains the only defensible verdict. The concrete test above would settle the issue as soon as the full text appears.","tokens_in":2101,"tokens_out":438,"duration_ms":13428,"concrete_test":"Once the full paper is available, extract the formal definition of decoding margin and the capacity lower-bound argument (expected in the main theorem or §3); recompute the leading constant of the isotropic capacity scaling and confirm it matches the known information-theoretic upper bound up to (1+o(1)); then check that the residual-stream insertion experiments use exactly the same closed-form weights without extra fine-tuning that would invalidate the optimality claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a closed-form MLP that simultaneously (i) matches information-theoretic capacity under isotropic embeddings and (ii) remains functional inside residual streams. The abstract presents “analyzing the decoding margin” as the single technical step that delivers both properties, yet supplies neither a definition of the margin, the precise conditions under which a positive margin guarantees correct fact retrieval, nor any sketch linking that margin to the claimed capacity lower bound. Without those, it is impossible to verify that the construction truly attains the IT optimum or that residual-stream and attention-geometry interactions do not introduce extra interference that would destroy the claimed scaling. The 10–104\times and 15–63\times parameter reductions are therefore unanchored numerical assertions whose validity collapses if the unspecified margin conditions fail for realistic embedding geometries.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript claims a closed-form construction of Transformer-compatible fact-storing MLPs that simultaneously (i) attains information-theoretically optimal fact-storage capacity scaling under isotropic embeddings, (ii) extends to arbitrary key/value geometries up to geometry-dependent penalization factors, and (iii) functions inside Transformer residual streams for factual recall. The stated technical lever is analysis of the MLP decoding margin (rather than storage capacity alone). Under isotropic embeddings the construction is reported to need 10–104× fewer parameters than prior constructions at matched fact count; inside Transformer blocks the reported reduction is 15–63×. A proof-of-concept modular fact-editing application (swapping an MLP) is also claimed.","tokens_in":2276,"tokens_out":701,"duration_ms":14356,"significance":"If the optimality, geometric generality, and residual-stream compatibility claims hold with the stated parameter reductions, the work would supply the first constructive account of near-optimal fact storage in Transformer MLPs, substantially improve on prior constructive baselines, and open a modular route to fact editing. Those outcomes would be of clear interest to both the theory and mechanistic-interpretability communities. The abstract-only submission, however, does not yet allow those claims to be verified.","major_comments":[{"comment":"The abstract presents “analyzing the decoding margin” as the single step that simultaneously yields IT-optimal capacity and Transformer compatibility, yet supplies neither a definition of the margin, the precise conditions under which a positive margin guarantees correct retrieval, nor a sketch linking that margin to a capacity lower bound. Without those, the central optimality claim cannot be checked.","section":null},{"comment":"The reported 10–104× (isotropic) and 15–63× (in-Transformer) parameter reductions are numerical assertions whose validity depends on the unspecified margin conditions, the choice of baselines, embedding dimension / width, and experimental protocol. None of these appear in the abstract, so the efficiency claims are unanchored.","section":null},{"comment":"Extension from isotropic embeddings to arbitrary geometries is stated only as “up to penalization factors depending on the embedding geometries.” The abstract does not define those factors, bound them, or state when they remain O(1). If the factors grow with the number of facts or with embedding anisotropy, the claimed optimal scaling does not transfer.","section":null},{"comment":"Transformer compatibility is asserted (“works inside Transformers,” “used within Transformer blocks for factual recall tasks at optimal capacity scaling”) without any statement of residual-stream or attention-geometry interference conditions. Residual addition and attention can introduce extra cross-fact interference that would invalidate a pure-MLP margin analysis; the abstract gives no argument that this does not occur.","section":null}],"minor_comments":[],"recommendation":"uncertain","confidential_remarks":"Only the abstract was available for review; the full manuscript (theorems, proofs, experimental protocols, tables) was not provided. Under those constraints a soundness verdict is impossible and the appropriate editorial action is to request the full paper (or to desk-reject for incompleteness). The stress-test concern about undefined decoding-margin conditions is well-founded on the abstract alone and should be the first item checked once the full text is in hand. I have no view on novelty relative to concurrent work without the full bibliography and related-work section."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing you need to know: this abstract claims the first closed-form MLP construction that stores facts at information-theoretically optimal scaling, handles arbitrary key/value geometries, and still works inside Transformer residual streams. If the full paper delivers, that is a real within-field advance for mechanistic interpretability and efficient fact packing.\n\nWhat looks new is the shift from pure storage capacity to decoding margin. Prior constructive work apparently left a gap between theory and the empirical packing rates people see in LLMs; the authors say analyzing the margin closes that gap and yields 10–104× parameter reductions at matched fact count (15–63× when the MLP sits inside a Transformer block). They also sketch a modular editing proof-of-concept by swapping the MLP. That package is cleanly framed against external IT bounds rather than self-referential fitting, which is good structure.\n\nThe soft spots are exactly what you expect from abstract-only material. The optimality claim under isotropic embeddings and the “same scaling up to geometry penalization factors” for arbitrary embeddings are numerical and geometric assertions we cannot verify. The stress-test note is right that “decoding margin” is never defined here, so we cannot check the conditions that supposedly guarantee both the capacity lower bound and residual-stream compatibility. Free parameters (geometry factors, width choices) remain opaque. None of that proves the claims false; it just means soundness is currently unscoreable.\n\nWho it is for: people who care about constructive models of factual recall, capacity allocation, and editing primitives. A serious referee should see the full theorems, proofs, and experiments. I would send it to peer review rather than desk-reject; the claimed combination of closed form, optimality, and Transformer compatibility is important enough to deserve that time. Bring it to reading group only after the PDF appears and someone has checked the margin analysis. I would not cite it yet.","headline":"Abstract-only claim of a closed-form, Transformer-compatible fact-storing MLP that hits IT-optimal capacity; interesting if true, but currently uncheckable.","tokens_in":2887,"tokens_out":481,"would_cite":false,"duration_ms":4266,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Closed-form MLPs store facts at optimal capacity inside Transformers, using far fewer parameters than prior constructions.","keywords":["fact storage","MLP","Transformer","information-theoretic capacity","decoding margin","closed-form construction","modular editing","key-value memory"],"falsifier":"Construct the claimed MLP for a known set of facts with isotropic embeddings of dimension d and verify whether its parameter count scales as Θ(n d) (optimal) while the decoding margin remains positive and factual recall accuracy stays high when the MLP is substituted into a Transformer block; a clear super-linear blow-up or collapse of the margin falsifies the claim.","tokens_in":2954,"feed_emoji":"🧠","tokens_out":765,"duration_ms":6013,"temperature":0.7,"pith_summary":"Language models store facts in their MLP layers, yet existing constructive accounts of that storage do not explain how real models reach information-theoretically optimal density. This paper supplies the missing construction: a closed-form MLP that is fully compatible with the residual stream and attention geometry of a Transformer, stores an arbitrary set of key–value facts, and matches the optimal capacity scaling under isotropic embeddings. The technical step is to analyze the decoding margin of the MLP rather than only its storage capacity; once that margin is controlled, the same construction continues to work for non-isotropic embeddings (up to geometry-dependent penalties) and can be dropped directly into Transformer blocks for factual recall. Empirically the resulting modules need one to two orders of magnitude fewer parameters than earlier fact-storing MLPs at the same number of facts, and they support modular editing by simple MLP swap.","feed_headline":"MLPs store facts at optimal density with 10–100× fewer weights","feed_subtitle":"Closed-form construction works inside Transformers and supports modular fact editing by MLP swap.","key_machinery":"Decoding-margin analysis of the MLP. By optimizing the geometric margin that separates correct from incorrect value embeddings after the MLP, rather than merely counting storage capacity, the authors obtain an explicit weight construction that simultaneously hits the information-theoretic optimum and stays compatible with residual-stream geometry.","core_discovery":"A closed-form, Transformer-compatible fact-storing MLP attains information-theoretically optimal fact-storage capacity scaling under isotropic embeddings, remains capacity-optimal (up to geometry penalties) for arbitrary embeddings, and can be inserted into Transformer blocks for factual recall while requiring 10–104× fewer parameters than prior constructions at matched fact count.","pith_inferences":["If the margin analysis generalizes, practitioners could pre-compute or surgically insert compact fact banks rather than relying solely on gradient-based memorization.","The same construction may supply a diagnostic: measure the effective decoding margin of a trained MLP to estimate how close it sits to the information-theoretic storage limit.","Geometry-dependent penalties suggest that embedding-regularization schemes that keep keys and values well-conditioned could further reduce the parameter overhead of fact storage."],"forward_implications":["Fact-storing MLPs can be written down in closed form and inserted into Transformers without destroying residual-stream geometry.","At matched fact count the construction needs 10–104× fewer parameters than earlier constructive baselines under isotropic embeddings, and 15–63× fewer when used inside Transformer blocks.","Arbitrary key/value geometries are handled with only a multiplicative capacity penalty that depends on the embedding condition numbers.","Modular fact editing becomes possible by simply swapping one fact-storing MLP for another inside a frozen Transformer."],"fun_headline_variants":["MLPs as Hebbians hit optimal fact storage density","Fact-storing MLPs need 10-104× fewer params than prior work","Closed-form MLP construction stores facts at optimal scaling","Transformer-ready MLPs attain info-theoretic fact capacity","Hebbian-style MLPs enable modular fact editing by swap"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That controlling the decoding margin of the MLP is enough to guarantee both information-theoretic optimality and seamless insertion into real Transformer residual streams for embeddings that are not perfectly isotropic.","fun_headline_variants_meta":{"raw":{"variants":["MLPs as Hebbians hit optimal fact storage density","Fact-storing MLPs need 10-104× fewer params than prior work","Closed-form MLP construction stores facts at optimal scaling","Transformer-ready MLPs attain info-theoretic fact capacity","Hebbian-style MLPs enable modular fact editing by swap"]},"model":"grok-4.5","effort":"low","cost_usd":0.004704,"raw_usage":{"total_tokens":1394,"prompt_tokens":819,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":47040000,"prompt_tokens_details":{"text_tokens":819,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":505,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":819,"tokens_out":70,"duration_ms":3988,"temperature":1.0,"reasoning_tokens":505,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T00:51:24.829168+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Construct the claimed MLP for a known set of facts with isotropic embeddings of dimension d and verify whether its parameter count scales as Θ(n d) (optimal) while the decoding margin remains positive and factual recall accuracy stays high when the MLP is substituted into a Transformer block; a clear super-linear blow-up or collapse of the margin falsifies the claim.","supporting_citations":[],"review_version":1}