{"id":"e5a3a43d-ba46-4df0-a2df-cea1d2babf00","arxiv_id":"2511.16362","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a kinetic-encoded quasi-2D self-assembly model, layer-nucleation events with only one bond create speed and encoding bottlenecks that can be removed by adding diagonal bonds to a small fraction of components, yielding Smax ~ N^{1-2/nc}.","lead":"This paper studies a toy model where protein-like tiles bind and unbind to a growing 2D structure, with speed and accuracy controlled by kinetic barriers rather than binding energies. It finds that a few rare, slow layer-starting events cause both slow and error-prone assembly, and shows that giving a small number of tiles extra binding connections removes both bottlenecks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Encoding-capacity scaling Smax∼N^{1−2/nc} relies on an untested independence count in SI Eq. (S11); a direct enumeration of confounding neighborhoods is needed.","rationale":"I considered the SI S3 stall regimes as an alternative; they qualify the main-text sentence 'µ>0, δ>δmin guarantee retrieval,' but the SI already presents them as separate regimes and the core connectivity mechanism still works in the retrieval window. I therefore agree with the Reader that Eq. (S11) is the weakest load-bearing point. The Reader already marked the verdict CONDITIONAL; my stress test does not move that verdict, so UNCHANGED. The paper's independent support (direct simulations of the z=4+ speed/accuracy improvement, scaling plots for τ_ret) is real and should be credited; the open question is whether the asymptotic encoding-capacity scaling is robust beyond the random-target mean-field estimate.","tokens_in":20652,"tokens_out":15517,"duration_ms":152244,"concrete_test":"Enumerate, for the same target constructions used in the paper, the number of incorrect species with ri≥nc for every critical (layer-nucleation) boundary neighborhood. Do this for N=ℓ^2 with ℓ=7,14,20,40 and S=2,4,8, both for uniformly random target permutations and for the z=4+, z=6+, z=8 designs. Average over at least 100 independent target samples and fit log NI vs log N at fixed S, and log NI vs log S at fixed N. If the slope in log N is -b with b≠nc−1, or the S-slope is a≠nc, Eq. (S11) fails. Also measure Smax directly via the sigmoid midpoint for ℓ=40 and 60 at z=6+; Eq. (6) predicts Smax≈c·N^{1/3}, so the ratio Smax(ℓ=60)/Smax(ℓ=40) should be (60/40)^{2/3}≈1.31; a ratio near 1 would indicate saturation and a failure of the claimed scaling.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative result is Eq. (6): Smax ∼ N^{1−2/nc}. Its derivation in SI S4 hinges entirely on Eq. (S11): NI ∼ (S−1)^nc / N^{nc−1}. This is a mean-field/random-target estimate: for each of the nc directions in a critical neighborhood, a wrong monomer has probability (S−1)/N of matching the correct encoded neighbor in some alternative target, and the directions are multiplied independently. Three features of the actual model can break this factorization: (i) within one target the encoded neighbors of a component are drawn without replacement, so the nc directions are not independent; (ii) a wrong monomer can satisfy different directions in different targets — the formula counts this, but the number of available targets is discrete (S−1), so for small S the independence approximation is poor unless the per-direction probabilities are uniform, which the deterministic extra diagonal bonds (z=4+, 6+) do not guarantee; (iii) the count is an average over critical events, whereas the low-error condition Nc perr≪1 requires the typical or worst-case location not to have larger NI. If the measured NI has an extra factor S or an N-exponent differing from nc−1, the Smax scaling and the predicted exponents 1/3 (z=6+) and 1/2 (z=8) shift. The numerical support in Fig. 6D/G uses ℓ=7,14,20 and a smoothing approximation nc≈z/2 (Eq. S18), so it cannot discriminate the asymptotic scaling from these alternatives. This is the most load-bearing point because it is the quantitative core of the \"encoding bottleneck\" claim; the z=4+ S=2 speed result is supported by direct simulation and would survive even if Eq. (6) were revised.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends a prior heteropolymer kinetic-encoding model to quasi-2D self-assembly. Components are square tiles with nearest-neighbor bonds and, in some designs, extra diagonal bonds; target structures are encoded not in binding energies but in the kinetics, via rates k^± = exp(r_i δ) and exp(r_i δ − µ), where r_i counts the number of neighbors consistent with any encoded target. Retrieval from a nucleation seed is studied analytically and by Gillespie simulation. The central claims are: (i) for nearest-neighbor connectivity z = 4, layer-nucleation events with a single bond (n_c = 1) create a speed bottleneck τ_ret ≈ exp(−δ) and, for S ≥ 2, an encoding bottleneck with S_max = 1; (ii) adding one diagonal bond per layer (z = 4+) raises n_c to 2, removes both bottlenecks, and gives an O(1) encoding capacity; (iii) for higher connectivity, S_max ∼ N^{1−2/n_c}, with explicit predictions S_max ∼ N^{1/3} at z = 6+ and ∼N^{1/2} at z = 8, summarized by S_max ∼ N^{1−4/z}. The paper also presents a compilation of assembly-factor counts and component connectivities for large protein complexes, and argues that kinetic control of a few critical binding events is a plausible design principle for fast, accurate assembly.","tokens_in":20967,"tokens_out":10375,"duration_ms":100496,"significance":"If the central derivation is sound, the paper offers a parameter-free, falsifiable scaling theory for a nontrivial kinetic-encoding problem in heterogeneous self-assembly. The main predictions—Eqs. (4)–(8)—have no fitted parameters and are checked against Gillespie simulations, which is a genuine strength. The idea that a small number of critical events, rather than bulk connectivity, controls both speed and accuracy is novel and potentially important for understanding assembly factors in ribosome and spliceosome assembly. The empirical Fig. 1 data provide useful biological motivation. However, the quantitative core rests on a combinatorial estimate in SI S4 that is not rigorously derived or directly tested, and the main-text sufficiency claim for the retrieval conditions is contradicted by the SI phase diagram. These issues are load-bearing for the paper's main claims, so the result is not yet established at the level claimed.","major_comments":[{"comment":"The main text states 'Conditions µ > 0 and δ > δmin result in accurate retrieval' (Sec. V.B). This is not correct as a general statement: SI Fig. S3C-D and Eq. (S4) show that a substantial part of the nominal region µ>0, δ>δmin is not in the transient-retrieval regime. In particular, above the blue line δ_blue ≈ µ/(r_dis − n_c) growth stalls near the seed size, and between the gray and blue lines the target is recovered without further growth. Thus Eq. (4) is only a lower bound on δ, not a sufficient condition. This is load-bearing because the paper's practical claim is that positive µ and sufficiently large δ guarantee fast, accurate retrieval. The main text should state the full retrieval conditions—including the upper bounds—or explicitly restrict the analysis to the µ≫δ regime used in the main simulations, and justify that the simulation parameters of Figs. 3–6 lie inside the retriev","section":"V.B and SI S3 (Eq. S4)"},{"comment":"The central encoding-capacity scaling S_max ∼ N^{1−2/n_c} (Eq. 6) is derived from the estimate N_I ∼ (S−1)^{n_c}/N^{n_c−1} in Eq. (S11). This estimate treats the n_c required neighbor matches as independent draws, allowing each match to come from a possibly different alternative target. In the actual model, targets are reshufflings of the same N species, the neighbors of a monomer within a single target are drawn without replacement, and a confounding monomer must be a valid component of a complete target. Correlations among the S target permutations are not controlled. Moreover, the condition N_c p_err ≪ 1 concerns the typical or worst-case critical event, while Eq. (S11) is an average estimate. Because Eqs. (6) and (8) are the quantitative core of the paper, this omitted justification is load-bearing. Please provide a derivation for a well-defined random-target ensemble, or directly me","section":"SI S4, Eq. (S11); Eq. (6)"},{"comment":"Equation (8), S_max ∼ N^{1−4/z}, is presented as a general scaling, but it is obtained by replacing the discrete n_c with the continuous approximation n_c ≈ z/2 (Eq. S18). Table II and Eq. (S15) show that n_c, N_c, and Ω_C jump at integer values of the number of extra bonds per layer, producing period-ℓ oscillations (Fig. S8). The fit in Fig. 6G uses only three system sizes and a continuous α = 1−4/z, so it does not establish Eq. (8) as a sharp asymptotic prediction. The paper should state explicitly that Eq. (8) is a coarse-grained interpolation, and use the discrete expressions (S15)–(S17) for quantitative comparisons, or provide a rigorous argument for why the fluctuations average out in the large-N limit.","section":"SI S5, Eq. (S18); Eq. (8)"}],"minor_comments":[{"comment":"The wording in Sec. VII 'To guarantee that n_c = 3, we consider assemblies with a bulk connectivity z = 6' is confusing because Table II lists n_c = 2 for z = 6. Please clarify that the design denoted z = 6+ (not z = 6 itself) ensures n_c = 3, or revise the sentence to avoid the apparent contradiction.","section":"VII and Table II"},{"comment":"The quantities N_c, Ω_C, and Ω_I are used in the main text before they are defined. Please define them at first use, or add a short table in Sec. V, since the reader otherwise has to go to SI S4 to understand Eq. (4).","section":"V"},{"comment":"The sentence 'We thus recall results from the polymerization study [35]' precedes Eq. (S11), which is not a literal result from that study—it is an extension to 2D neighborhoods with n_c ≥ 2. Please present the derivation explicitly rather than recalling it.","section":"SI S4"},{"comment":"The inset claims S_max shows negligible N-dependence for n_c = 2, but the figure uses only ℓ = 7, 14, 20. Please add more system sizes or error bars to make the O(1) claim visually convincing.","section":"Fig. 5D"},{"comment":"The term 'target lifetime' is defined in the main text as the time to add a few incorrect monomers, while SI S3 uses a related but different notion (time to add an extra monomer chunk). Please align the definitions or explicitly distinguish the two quantities.","section":"SI S3"}],"recommendation":"major_revision","confidential_remarks":"The conditional verdict of the reader is appropriate. The most serious issues are (1) the SI S3 phase diagram contradicts the main-text sufficiency claim for Eq. (4), and (2) the S_max scaling rests on the insufficiently justified independence count in Eq. (S11). Both are fixable in revision: the first by stating the full retrieval region, the second by a direct enumeration or a rigorous combinatorial derivation. I would not reject the paper, since the central mechanism is plausible and the parameter-free predictions are a strength, but the quantitative claims should not be published in their current form without addressing these points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this paper. First, the core mechanism is clean: in a quasi-2D kinetic-encoding model, the slowest events—nucleating a new layer with a single bond (nc=1)—are also the events where incorrect monomers compete on equal footing, so speed and accuracy fail together. Adding one diagonal bond per layer (z=4+, only 2ℓ tiles affected) raises nc to 2 and eliminates both bottlenecks, with retrieval time dropping from ℓ exp(−δ) to ℓ² exp(−2δ). That result is supported by Gillespie simulations and is the main reason to read the paper.\n\nSecond, the encoding-capacity scaling Smax ~ N^{1−2/nc} is the quantitative centerpiece of the \"encoding bottleneck\" claim, and it rests on a single heuristic count, Eq. (S11), which treats confounding neighborhoods as independent random reshufflings. The stress-test note is right that this can break: within a target the encoded neighbors are drawn without replacement, the per-direction match probabilities aren't uniform when diagonal bonds are added deterministically, and the count is an average while the low-error condition requires typical or worst-case behavior. The numerical check against ℓ=7,14,20 fits an exponent but can't discriminate the asymptotic scaling from nearby alternatives. So the capacity law is plausible but not yet demonstrated. The z=4+, S=2 speed result does not depend on this count and stands on its own.\n\nA second, smaller issue: the main text says \"µ>0 and δ>δmin result in accurate retrieval,\" but SI S3 shows a substantial stall region inside exactly those parameters. The SI is honest about it; the main text overstates.\n\nOtherwise the paper is clean: no fitted parameters, scalings derived and compared to stochastic simulation, and the biological motivation is framed as a hypothesis rather than a conclusion. No code or data shipped, so reproduction requires reimplementation, but the model is simple enough that this is a moderate inconvenience.\n\nWho gets value? People working on kinetic proofreading in assembly, multifarious structures, or DNA tile design will find the layer-nucleation bottleneck idea useful and design-relevant. The capacity scaling should be treated as an open conjecture until Eq. (S11) is stress-tested.\n\nRecommendation: send it to peer review. Ask the authors to either tighten the derivation of Smax (direct enumeration of confounding neighborhoods in small systems, or a provable bound) or explicitly frame Eq. (6) as heuristic. And reconcile the main-text retrieval condition with the stall regimes. With those revisions, the core mechanism is worth publishing on its own.","headline":"A kinetic-encoding model with a genuinely nice bottleneck mechanism; the central capacity scaling rests on a heuristic count that needs shoring up before the strong claims are accepted.","tokens_in":21522,"tokens_out":2695,"would_cite":true,"duration_ms":27991,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single extra bond per layer resolves both the speed and encoding bottlenecks of quasi-2D heteromeric self-assembly, by raising the critical layer-nucleation events from one to two bonds and yielding an encoding capacity that scales as N^{","keywords":["kinetic encoding","self-assembly","heteromeric complexes","assembly factors","speed-accuracy tradeoff","encoding capacity","connectivity","bottlenecks"],"falsifier":"Measure the maximum number of codable structures Smax for a fixed system size N under controlled target sets: with targets engineered to share correlated motifs (e.g., repeated sub-neighborhoods), the capacity should fall below the predicted N^{1-2/nc}; conversely, with random reshufflings it should match. A second test: for z=4+, the retrieval time should scale as τret ~ N exp(-2δ); if the measured δ-dependence instead remains exp(-δ), the extra bond has not actually raised the critical bond number, falsifying the bottleneck-suppression mechanism.","tokens_in":20490,"feed_emoji":"🧩","tokens_out":5772,"duration_ms":56707,"temperature":0.7,"pith_summary":"This paper studies a simplified model of how large multi-protein complexes assemble quickly and reliably from a crowded mixture. The authors encode target structures purely in reaction kinetics, not binding energies: monomers bind faster when they match a local neighborhood of a target. They show that in a two-dimensional lattice model, the slowest growth step—nucleating a new layer with only one bond—also becomes the step where wrong components cannot be distinguished, so speed and accuracy fail together. Their central result is that adding a single extra bond per layer, so that layer nucleation creates two bonds instead of one, removes both bottlenecks at once, with negligible change to average connectivity. They further show that increasing connectivity to three bonds per critical event raises the number of structures that can be encoded simultaneously from a constant to a power of the system size.","feed_headline":"One extra bond per layer unlocks fast, accurate assembly","feed_subtitle":"Raising layer-nucleation bonds from one to two removes speed and encoding bottlenecks in a kinetic model of protein self-assembly.","key_machinery":"The load-bearing construct is the critical addition event: the growth step with the smallest number of bonds n_c, defined through the mini-max rate k_c = min_t max_{i,x} exp(r_i(N_x) δ). In the irreversible high-discrimination regime, n_c controls three outputs: the retrieval time τ_ret ≈ (N_c/Ω_C) exp(-n_c δ), the discrimination threshold δ_min = (1/n_c) ln(N_c Ω_I/Ω_C), and the encoding capacity S_max ~ N^{1-2/n_c}. The capacity scaling follows from a combinatorial count of 'confounding' monomers that share n_c correct partners with the intended one, estimated as N_I ~ (S-1)^{n_c} / N^{n_c-1} under the assumption that target rearrangements are random. Increasing local connectivity—one extr","core_discovery":"The paper's central claim is that, in quasi-2D heteromeric self-assembly with kinetic encoding, the critical events that limit assembly speed are the same events that limit encoding accuracy: the nucleation of new layers, where a monomer attaches with a single bond (nc = 1). With nearest-neighbor-only connectivity (z = 4), these events slow retrieval time to τret ≈ exp(-δ) and cap the number of simultaneously encodable structures at Smax = 1. Adding one diagonal bond per layer (z = 4+, only 2ℓ of N components affected) lifts the critical bond number to nc = 2, making retrieval exponentially faster, τret ≈ N exp(-2δ), and eliminating combinatorial errors for a few targets. More generally, for","pith_inferences":["The bottleneck identity found here—one-bond events being both the slowest and the least discriminating—may generalize beyond square lattices: in any growth process where the lowest-bond-count step has a combinatorial ambiguity, speed and accuracy will fail together, and adding a bond at that step should decouple them.","A testable extension of the paper's logic is to three-dimensional assembly: the critical sites would be nucleation events on faces or edges, and the scaling of S_max with N might change because the surface-to-volume ratio differs; this could connect to experimentally observed assembly-factor distributions in 3D complexes.","The random-reshuffling counting in SI Sec. S4 implies that the scaling Smax ~ N^{1-2/nc} should be robust to moderate target overlap, but would break for correlated target sets—this could be probed by designing structures with shared motifs and measuring whether capacity degrades faster than power-law.","The paper's kinetic-encoding framework, applied to two structures S=2, predicts a ~50% error plateau at large δ for z=4 that is insensitive to δ; this is a sharp, experimentally falsifiable signature distinguishing kinetic from energetic encoding."],"forward_implications":["If correct, the model predicts an exponential speedup in assembly time (from exp(-δ) to exp(-2δ) or exp(-3δ)) from a connectivity change affecting only O(√N) of N components.","It predicts that encoding capacity scales as a power of system size once critical events create at least two bonds: S_max ~ N^{1/3} when three bonds per critical event are enforced.","It identifies combinatorial (not thermal) errors as the fundamental limit on storing multiple structures with shared components, independent of the discrimination energy δ.","The mechanism suggests why assembly-factor counts grow with complex size: factors effectively supply local connectivity at critical sites, a role consistent with observed data on ribosomes and other large complexes.","It suggests design rules for synthetic self-assembling systems (e.g., DNA tiles): selectively boosting the connectivity of a few components should increase both yield and speed without altering the majority of interactions."],"fun_headline_variants":["A single extra bond per layer unlocks fast, error-free self-assembly","One more bond per layer removes speed and encoding bottlenecks","Small connectivity boost resolves protein assembly bottlenecks","Raising critical bond count from one to two accelerates assembly","Kinetically controlling key binding events beats energy specificity"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central scaling law Smax ~ N^{1-2/nc} rests on treating the set of target structures as independent random reshufflings when counting how many wrong monomers share n_c bonds with the correct one (SI Sec. S4, Eq. S11); if real targets are correlated in their neighborhoods, the predicted capacity scaling would not hold.","fun_headline_variants_meta":{"raw":{"variants":["A single extra bond per layer unlocks fast, error-free self-assembly","One more bond per layer removes speed and encoding bottlenecks","Small connectivity boost resolves protein assembly bottlenecks","Raising critical bond count from one to two accelerates assembly","Kinetically controlling key binding events beats energy specificity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000256,"raw_usage":{"total_tokens":1401,"prompt_tokens":721,"completion_tokens":680,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":465,"completion_tokens_details":{"reasoning_tokens":601}},"tokens_in":465,"tokens_out":680,"duration_ms":7460,"temperature":1.0,"reasoning_tokens":601,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T21:07:35.140993+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the maximum number of codable structures Smax for a fixed system size N under controlled target sets: with targets engineered to share correlated motifs (e.g., repeated sub-neighborhoods), the capacity should fall below the predicted N^{1-2/nc}; conversely, with random reshufflings it should match. A second test: for z=4+, the retrieval time should scale as τret ~ N exp(-2δ); if the measured δ-dependence instead remains exp(-δ), the extra bond has not actually raised the critical bond number, falsifying the bottleneck-suppression mechanism.","supporting_citations":[],"review_version":1}