{"id":"4bed8bf4-ad42-4b2f-9398-7616c5003d42","arxiv_id":"2506.14006","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Simulated directed evolution can train competitive dimerization networks to perform multiclass classification with performance comparable to gradient descent.","lead":"This paper shows, in computer simulations, that networks of molecules which bind into dimers can be trained by simulated evolution to classify input patterns. If the approach transfers to the lab, it would be a step toward computing without digital electronics, using chemistry instead.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that CDNs can be trained in vitro rests on Eq. 3's assumption that each Kij mutates independently; real sequence mutations correlate all affinities of a species, so the evolutionary search may exploit an unrealistically large parameter space.","rationale":"We read the paper as a simulation-based proposal that CDNs can be trained by directed evolution. The numerical demonstration is transparent and the comparison with gradient descent is informative. The weakest point is Eq. 3: it grants the evolutionary process independent control over each Kij, which is not achievable when Kij derives from the sequences of two molecules. This is a genuine correctness risk for the central claim, because the reported success may depend on this extra tunability. The reader's weakest_assumption identifies exactly this issue, so we agree. However, the paper explicitly acknowledges the simplification and the protocol is in silico in the first place; the concern does not invalidate the conceptual contribution but does require an additional test to support the in vitro claim. Hence the conditional acceptance stands unchanged. A concrete test with correlated mutations would settle whether the assumption is load-bearing.","tokens_in":10888,"tokens_out":8179,"duration_ms":94891,"concrete_test":"Repeat the evolutionary training protocol from SI Appendix C on the same three-class cocktail task, but replace Eq. 3 with a correlated mutation model: for each round, draw a per-species Gaussian noise u_i ~ N(0, σ_mut^2) and update ln K_ij -> ln K_ij + u_i + u_j (with a small independent component to preserve some variability). This enforces the physical constraint that a mutation in species i shifts all of its binding affinities together. Compare the final median mutual information and on/off contrast against Figs. 3i-j. If performance drops substantially below the reported near-maximal values, the independence assumption in Eq. 3 is the reason for the demonstrated learning, and the in vitro claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. 3 models directed evolution as an independent multiplicative random walk of every association constant Kij. In the proposed experimental implementation, however, Kij is determined by the sequences of species i and j; a mutation in one master sequence changes the binding free energy of that species with all its partners simultaneously, introducing correlations that Eq. 3 does not capture. The training algorithm thus effectively optimizes O(N^2) independent parameters, while a sequence-based realization offers at most O(N) sequence degrees of freedom (plus per-species concentrations). This inflates the search space and may make the reported near-maximal mutual information and contrast (Figs. 2-3) an artifact of unrealistic mutational freedom. The authors acknowledge the simplification in the sentence about a 'more detailed study' but provide no test of whether correlated mutations would still permit learning. Since the abstract claims in vitro trainability, this unvalidated assumption is load-bearing: if real mutations cannot independently tune individual pairs, the evolutionary search may fail in practice.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a chemical learning framework based on Competitive Dimerization Networks (CDNs), in which reversible dimerization equilibria of N molecular species implement a classifier from input concentrations to output monomer fugacities. A directed-evolution training protocol is outlined (Fig. 1c), with mutations modeled as independent biased multiplicative random walks on each association constant Kij (Eq. 3). The authors simulate this protocol numerically for three-class classification, reporting high on/off contrast and mutual information, robustness to input noise, sparsity induced by a drift parameter, and performance comparable to gradient descent (Figs. 2–5). Code and data are provided.","tokens_in":11077,"tokens_out":3808,"duration_ms":43485,"significance":"The CDN classifier is an intriguing physical-computation concept: the mass-action equilibrium equations are standard, and the mapping from chemical concentrations and affinities to a nonlinear classifier is clearly presented. The held-out test cocktails used for evaluation prevent the loss functions from being self-fulfilling, and the comparison with gradient descent is a useful sanity check. The availability of code and an interactive figure is a strength. However, the manuscript's central claim that CDNs 'can be trained in vitro' is not supported by experiments, and the mutation model may unrealistically inflate the searchable parameter space. If the proposed protocol can be realized and the mutation assumption can be relaxed, this would be a notable advance; at present the evidence is purely in silico.","major_comments":[{"comment":"The abstract states that CDNs 'can be trained in vitro through directed evolution' and the Discussion says the study 'demonstrated their in vitro training through directed evolution.' However, no wet-lab experiment is reported anywhere in the manuscript. The training protocol in Fig. 1c is a proposal, and the results in Figs. 2–5 and the SI are entirely numerical. Moreover, the protocol text says variants are 'ranked in silico' by the loss function, so even the proposed selection step is computational. Please revise the claims to 'proposed' or 'demonstrated in silico,' and clearly separate the proposal of an experimental protocol from the actual numerical demonstration.","section":"Abstract; Significance Statement; Discussion"},{"comment":"Eq. 3 models each Kij as an independent biased multiplicative random walk. In any sequence-based implementation (DNA oligomers as described), a single mutation in a master sequence changes the binding free energy of that species with all of its partners simultaneously, inducing strong correlations among the Kij. The simulations therefore effectively optimize O(N^2) independent parameters, whereas a sequence realization offers only O(N) sequence degrees of freedom plus concentrations. The near-optimal mutual information and contrast in Figs. 2–3 may rely on this unrealistically large search space. The authors acknowledge that a 'more detailed study' would compute Kij from sequences (citing Ref. 27), but they do not test whether correlated mutations still permit learning. Since the in vitro trainability claim depends on this assumption, please add a test with correlated mutations, for example a sequence-based model in which Kij is derived from pairwise sequence complementarity, or a block-correlated mutation model that perturbs all affinities of a mutated species jointly.","section":"Eq. 3; 'In-vitro Training via Directed Evolution'"},{"comment":"The training protocol includes mutating both Kij and the hidden/output concentrations, with 150 variants per generation. The reported performance is averaged over multiple drift (EL) and regularization (GD) parameters, but the manuscript does not report the variance or error bars for the mutual-information and contrast values in Figs. 3e–j and 5a–b. Since the loss landscape is described as 'highly degenerate and rugged,' a small number of runs (e.g., 20 in Fig. 5c–d) may not establish that EL reliably matches GD. Please provide error bars or a statement of the number of independent realizations used in each heatmap/scatterplot.","section":"Results; SI Appendix C"}],"minor_comments":[{"comment":"The two displayed forms of the contrast loss are not obviously equivalent; please define the averages and the indicator variable phi more explicitly, and show the algebra connecting the two expressions.","section":"Eq. 4"},{"comment":"The derivative of the contrast loss appears to have index inconsistencies (e.g., delta_ij with a beta index and a j = argmax condition). Please check and clarify.","section":"SI Appendix, Eq. 21"},{"comment":"The caption states that 'all displayed interactions have identical association constants, fixed at the upper saturation limit,' while the main text says edge thickness reflects log Kij. If all Kij are identical, the thicknesses should be uniform; please resolve this discrepancy.","section":"Fig. 2e caption"},{"comment":"The term 'fugacity' is used for the fraction of total concentration that remains unbound, which is not the standard thermodynamic fugacity; consider using 'free fraction' or 'unbound fraction' to avoid confusion.","section":"Model and Protocols"},{"comment":"The drift parameter is called 'negative drift' in the main text but eta0 appears positive; please define the sign convention consistently.","section":"SI Appendix C"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to attract attention because of the in vitro claim, which currently overstates the evidence. The mutation-model issue (Eq. 3) is the most substantive technical concern and should be addressed with an additional simulation, not just a caveat. The comparison with gradient descent is interesting but needs error bars."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The new thing here is that a competitive dimerization network can be trained by a simulated directed-evolution protocol to do multiclass classification, and that a contrast-based loss function buys you better on/off separation than MSE. The comparison with gradient descent is thorough and honest: GD is slightly better, and the two methods produce wildly different network structures with similar performance. The code is on GitHub, and the SI has a clean recursive scheme for exact gradients. For an in silico proof of concept, I think the central numerical result holds up.\n\nThe soft spot is the gap between the abstract and the evidence. The abstract and significance statement say the networks 'can be trained in vitro', but there is no experiment here—everything is simulation. That overclaim needs to be pulled back to 'can be trained in silico with a plausible wet-lab protocol.'\n\nThe bigger scientific concern is the mutation model. Equation (3) treats each Kij as an independent multiplicative random walk. In the proposed implementation, mutations act on the sequences of master sequences, so one mutation changes the affinity of a species to all of its partners simultaneously. The optimizer is effectively tuning O(N^2) independent parameters, while a sequence-realization has far fewer degrees of freedom. The authors acknowledge this in a sentence and point to a 'more detailed study', but they never test whether correlated mutations still allow learning. That is load-bearing, because the paper's main claim is that in vitro training works.\n\nMinor points: the performance metrics have no error bars, and the network sizes and hyperparameters look like they were chosen by convenience rather than by any principle. The sparsity-vs-drift observation is nice, but not central.\n\nWho is this for: people working on physical learning, DNA/protein computing, and reservoir computing. It deserves a serious referee. I'd recommend returning it with the request to (1) fix the in vitro language, (2) add a simulation with correlated mutations—even a simple toy model where a single mutation shifts all affinities of a species by a correlated amount—and (3) add error bars or at least confidence intervals on the MI and contrast.\n\nMy own verdict: the paper is a solid in silico contribution, but the headline claim needs to be scaled back. Send it to review.","headline":"Simulated directed evolution trains dimerization networks for classification, but the in vitro claim outruns the evidence and the independent-mutation model is the soft underbelly.","tokens_in":11602,"tokens_out":2966,"would_cite":true,"duration_ms":29859,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Competitive dimerization networks can be trained by directed evolution to classify noisy input mixtures, with performance rivaling gradient descent.","keywords":["chemical learning","competitive dimerization networks","directed evolution","molecular computation","mass-action kinetics","physical learning","multiclass classification","mutual information"],"falsifier":"A direct experimental test would be to run the proposed in vitro directed evolution protocol on a three-class, six-input CDN made of non-complementary DNA oligomers and measure classification accuracy and mutual information at 20% input noise; if the trained chemical classifier does not approach the simulated near-maximal mutual information, the central claim is falsified. A computational falsifier: replace the independent multiplicative random walk of Eq. (3) with correlated mutations affecting multiple $K_{ij}$ (for example, via shared sequence segments) and check whether the evolved network still reaches the same performance; if performance collapses, the independent-walk assumption is load-bearing.","tokens_in":10654,"feed_emoji":"🧬","tokens_out":10500,"duration_ms":89816,"temperature":0.7,"pith_summary":"This paper tries to establish that a mixture of molecules that reversibly bind into dimers — a competitive dimerization network — can serve as a trainable physical computer. The authors show that by subjecting the network to rounds of mutation, selection, and amplification, it can learn to map noisy input cocktails to one-hot encoded output classes, with each molecular species acting like a neuron and each binding affinity acting like a synaptic weight. They report near-maximal mutual information between inputs and outputs, on/off contrast spanning orders of magnitude, and performance closely correlated with in silico gradient descent training. If true, this would provide an in vitro-trainable, energy-efficient platform for analog molecular computation that needs no digital hardware or explicit parameter tuning.","feed_headline":"Evolution trains chemical dimer networks to classify inputs","feed_subtitle":"Mutation, selection, and amplification train a DNA-based mixture to classify noisy inputs without digital hardware.","key_machinery":"The central objects are the equilibrium equations of the dimerization network: for total concentrations $c_i$ and free concentrations $x_i = f_i c_i$, mass action gives $c_i = x_i + \\sum_j K_{ij} x_i x_j$, and the output fugacities solve $f_i = 1/(1 + \\sum_j K_{ij} c_j f_j)$. This fixed-point system is what maps an input cocktail (the input concentrations $c_{\\mathrm{in}}$) to output fugacities $f_{\\mathrm{out}}$. Training is driven by the mutation model of Eq. (3), where each association constant undergoes a biased multiplicative random walk, $K_{ij}(t+1) = K_{ij}(t) e^{\\sigma_{\\mathrm{mut}}(\\eta_{ij}(t)-\\eta_0)}$, with drift $\\eta_0$ toward weaker binding acting as a regulariser; selection uses the contrast loss of Eq. (4), which penalises the worst output component's ratio of mean 'off' fugacity to geometric-mean 'on' fugacity. The same equilibrium equations admit linear-response gradients (SI section D), enabling a direct comparison with gradient descent training.","core_discovery":"The central discovery is that the equilibrium state of a competitive dimerization network — described by the mass-action equations $c_i = x_i + \\sum_j K_{ij} x_i x_j$ with output fugacities $f_i = 1/(1 + \\sum_j K_{ij} c_j f_j)$ — is expressive enough to act as a nonlinear classifier, and that its parameters (the association constants $K_{ij}$ and the hidden-layer concentrations) can be tuned by an evolutionary protocol rather than by gradient descent. Using a contrast-enhancing loss function, the evolutionary search produces output fugacity distributions whose 'on' and 'off' states are separated by several orders of magnitude and whose mutual information approaches the theoretical maximum of 1. The same model trained by gradient descent gives highly correlated performance, despite the network being bidirectional, having strictly nonnegative weights, and using the mass-action activation function $1/(1+y)$.","pith_inferences":["The paper only demonstrates networks with $n_{\\mathrm{out}} = 3$ output classes, so a natural untested extension is whether the evolutionary search remains effective for larger class counts, where the loss landscape may be more rugged.","The contrast loss uses a geometric mean of 'on' fugacities, which suggests the trained networks may be insensitive to the absolute concentration scale of the inputs; if so, the protocol would be especially convenient for biosensing applications where absolute concentrations are hard to control.","The authors draw an analogy to Hopfield networks but do not explore it; if the analogy holds, the same CDN framework could implement associative memory, storing multiple input patterns as attractors rather than only classifying them.","Because the mutation model multiplies each $K_{ij}$ independently, the training dynamics depend only on ratios of binding affinities; this suggests the protocol's behaviour is scale-invariant, which could simplify experimental calibration."],"forward_implications":["Once trained, the chemical classifier requires no digital computation: a new input cocktail directly produces output fugacities (read out, for example, via FRET) that identify the class.","The training protocol is compatible with existing DNA-based CDN implementations, using PCR-based random mutagenesis, restriction cleavage, and dilution, so the framework is experimentally realisable with current molecular biology techniques.","The drift parameter $\\eta_0$ acts as a regulariser that produces sparser networks with fewer strong interactions, making the trained dimerization networks more interpretable.","Because evolutionary learning and gradient descent give closely correlated performance, an in silico design phase can be used to pre-select association constants, which are then refined in vitro by directed evolution in a hybrid training strategy.","The loss function matters for output dynamic range: the contrast loss gives far better on/off separation than mean-squared error, while mutual information saturates near its maximum for both, so the choice of loss governs whether the goal is clean thresholding or information transmission."],"supporting_citations":[{"why":"Establishes the mathematical treatment of reversible protein-binding networks that the CDN model is built on.","marker":"[17]"},{"why":"Describes how perturbations spread in reversible reaction networks, framing the input-output response used here.","marker":"[18]"},{"why":"Demonstrates the computational capabilities of many-to-many protein interaction networks, motivating the CDN approach.","marker":"[19]"},{"why":"Shows contextual computation by competitive protein dimerization networks, the direct biological analogue of the proposed classifier.","marker":"[20]"},{"why":"Provides the experimental demonstration of DNA-based CDNs translating multiple inputs into programmed outputs, the substrate for the in vitro protocol.","marker":"[21]"},{"why":"Introduces non-complementary DNA oligomer systems that allow CDN implementation with sequences whose affinity can be mutated.","marker":"[22]"},{"why":"The Hopfield network analogy connects the equilibrium state of the CDN to associative memory and collective computation.","marker":"[25]"},{"why":"Extends the analogy to graded-response neurons, matching the fugacity outputs of the chemical network.","marker":"[26]"},{"why":"Supplies a sequence-to-affinity model that a more detailed treatment would use to compute $K_{ij}$ from actual DNA mutations.","marker":"[27]"}],"fun_headline_variants":["Evolution trains molecular networks for chemical classification","DNA dimers learn classification via directed evolution","Chemical networks evolve to classify without digital tech","Dimerization networks learn through evolutionary chemistry"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole training protocol rests on the assumption that each round of mutagenesis changes every binding affinity independently, as a multiplicative random walk; if real mutations move several affinities together, the evolutionary search may behave very differently from the simulations.","fun_headline_variants_meta":{"raw":{"variants":["Evolution trains molecular networks for chemical classification","DNA dimers learn classification via directed evolution","Chemical networks evolve to classify without digital tech","Dimerization networks learn through evolutionary chemistry"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1363,"prompt_tokens":892,"completion_tokens":471,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":426}},"tokens_in":508,"tokens_out":471,"duration_ms":5312,"temperature":1.0,"reasoning_tokens":426,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:24:28.488592+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct experimental test would be to run the proposed in vitro directed evolution protocol on a three-class, six-input CDN made of non-complementary DNA oligomers and measure classification accuracy and mutual information at 20% input noise; if the trained chemical classifier does not approach the simulated near-maximal mutual information, the central claim is falsified. A computational falsifier: replace the independent multiplicative random walk of Eq. (3) with correlated mutations affecting multiple $K_{ij}$ (for example, via shared sequence segments) and check whether the evolved network still reaches the same performance; if performance collapses, the independent-walk assumption is load-bearing.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the mathematical treatment of reversible protein-binding networks that the CDN model is built on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes how perturbations spread in reversible reaction networks, framing the input-output response used here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates the computational capabilities of many-to-many protein interaction networks, motivating the CDN approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows contextual computation by competitive protein dimerization networks, the direct biological analogue of the proposed classifier."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the experimental demonstration of DNA-based CDNs translating multiple inputs into programmed outputs, the substrate for the in vitro protocol."},{"cited_title":"Chem.15, 70–82 (2023)","cited_arxiv_id":null,"evidence_quote":"Introduces non-complementary DNA oligomer systems that allow CDN implementation with sequences whose affinity can be mutated."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Hopfield network analogy connects the equilibrium state of the CDN to associative memory and collective computation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends the analogy to graded-response neurons, matching the fugacity outputs of the chemical network."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies a sequence-to-affinity model that a more detailed treatment would use to compute $K_{ij}$ from actual DNA mutations."}],"review_version":1}