{"id":"c2a7d77f-a4e3-4759-874d-d17889d24025","arxiv_id":"2505.08531","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A building-block-aware diffusion model generates novel, large-unit-cell MOF crystal structures, and one model-suggested MOF was synthesized with a structure close to, but not identical to, the prediction.","lead":"This paper trains an AI model to design new metal-organic framework (MOF) crystals by learning and recombining their building blocks: metal nodes, organic linkers, and connection patterns. One MOF suggested by the model was made in the lab, though the final crystal differed slightly from the AI's blueprint.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validity metric is inherited from the curation cutoffs; the reported 52% validity may overstate synthesizability because edge angle <140° and node RMSD <0.3 Å are treated as invalid by construction.","rationale":"The reader's weakest_assumption is essentially identical: the validity metric inherits the curation cutoffs (0.3 Å RMSD and 140-degree edge angle), and Section 2b explicitly admits that edges below 140 degrees can form valid MOFs. The reader's verdict is CONDITIONAL, and my concern does not move that verdict; it sharpens the reason for conditionality. My concrete test is a direct calculation that would settle whether the concern lands: recompute validity after removing the curation cutoffs from the validity definition. The test is feasible because the authors have released the training data and the predecessor code, and the generated MOF set is described as 9,712 structures with stated validity, novelty, and uniqueness statistics. The experimental synthesis is a meaningful existence proof, but it does not calibrate the bulk validity metric, and the paper itself honestly notes the synthesized structure differs from the model prediction. I also note the code link points to OA-ReactDiff, not a BBA-specific implementation, which is a reproducibility concern but secondary to the validity-metric issue. I agree with the reader that the abstract overclaims 'great geometric validity' relative to the actual evidence; the concrete test would determine whether this is merely an overstatement or a genuine systematic overcount.","tokens_in":18410,"tokens_out":2259,"duration_ms":18910,"concrete_test":"Re-evaluate the 9,712 generated MOFs with a validity definition that does not reuse the curation cutoffs: (1) count as invalid only structures that fail a genuine sanity check (unphysical connectivity, floating atoms, impossible bond lengths/angles, severe steric clashes), and (2) separately relax the edge angle threshold to <140 degrees and the node RMSD threshold to >0.3 Å, applying the same sanity-check-only criterion to all structures regardless of angle or RMSD. If the sanity-check-only validity remains near 52%, the concern is alleviated. If it rises sharply, the reported validity metric conflates curation filters with genuine geometric quality.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that BBA MOF Diffusion samples MOFs with 'great geometric validity' rests on a validity definition that is inherited from the training-set curation thresholds. In Section 2b, a generated MOF is valid only if all building blocks are valid, and edge validity is defined by the same 140-degree angle cutoff used to curate the training data. The paper itself acknowledges (Section 2b) that CoRE MOF contains edges with angles below 140 degrees that form valid MOFs, with sulfonyldibenzene as an explicit example. Similarly, node validity is defined by an RMSD threshold of 0.3 Å relative to idealized net positions — the identical criterion used to select training examples. Thus the 52% overall validity, the 63% edge validity, and the 82% node validity are not independent measures of geometric plausibility; they are, to a substantial degree, a restatement of the curation filters. A generated structure that is actually synthesizable but has an edge angle of 139 degrees is counted as invalid, while a structure that passes the angle filter can still be synthetically inaccessible. The abstract's phrase 'great geometric validity' is therefore not supported by the reported metrics in the sense that the reader of the abstract would naturally understand it. The experimental synthesis provides an honest external anchor, but it does not calibrate the validity metric: the synthesized MOF differs from the model prediction in solvent coordination, and the paper does not report how the model's predicted structure would have scored on the validity metric. This concern is load-bearing because the main quantitative evidence for 'great geometric validity' is the 52% validity number, and that number is partially a tautology with the curation thresholds rather than an independent assessment of synthesizability or structural soundness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces BBA MOF Diffusion, an SE(3)-equivariant diffusion model that generates MOF crystals from a building-block representation consisting of an inorganic node, an organic edge, and a topological net, with the joint distribution learned from CoRE MOF 2019. Building blocks are denoised independently and assembled with PORMAKE on four nets (dia, nbo, sra, pcu). The authors report 63% edge validity, 82% node validity, and 52% overall geometric validity among 9,712 generated MOFs, unit cells up to 904 atoms, 27% novelty (13% novel and unique), conditional edge generation via inpainting, and an experimental synthesis of [Zn(1,4-TDC)(EtOH)2] characterized by PXRD, TGA, and N2 sorption. The paper candidly discusses its limitations: only single-node, single-edge MOFs and four nets are treated, property-guided generation is not yet implemented, and chemical compositions are assumed known.","tokens_in":18625,"tokens_out":10729,"duration_ms":103517,"significance":"If the claims hold, the paper represents a meaningful advance: representing MOFs as jointly diffused building blocks decouples the all-atom graph size handled by the scoring network from the unit-cell atom count, scaling generation to roughly 900 atoms, and the model can produce novel nodes and edges rather than recombining known fragments. The open code and data availability, the explicit statement of the four-nets limitation, and the attempt at external experimental validation with PXRD, TGA, and BET data are concrete strengths, and an experimental anchor of this kind is rare in generative-model papers. The significance is contingent on whether the reported validity numbers mean anything beyond the curation thresholds used to construct the training set, which is the central concern raised below.","major_comments":[{"comment":"The headline validity numbers are not independent of the curation procedure. Edge validity is defined by the same 140–180 degree connecting-point angle criterion used to curate the training set, and node validity is defined by the same <0.3 Å net-compatibility RMSD criterion described in Section 4b. The paper itself notes in Section 2b that edges with angles below 140 degrees (e.g., sulfonyldibenzene) occur in CoRE MOF and can form valid MOFs, so the 63%/82%/52% figures conflate 'compatible with the curation heuristics' with 'geometrically plausible and synthesizable'. Passing the filters does not imply a synthesizable structure, and failing them does not imply an invalid one; the abstract's phrase 'great geometric validity' is therefore not supported as an independent measure of physical plausibility. I recommend re-evaluating a subset of generated MOFs with an external structure-sanity check (for example, local coordination-environment validation or geometry relaxation) and/or reporting the same validity statistics for a non-generative baseline so the reader can calibrate the 52% figure.","section":"Section 2b / Section 4b"},{"comment":"The abstract states that PXRD, TGA, and N2 sorption 'confirm its structural fidelity', but Section 2d reports that the synthesized structure differs from the model prediction: two TDC ligands predicted to coordinate to Zn are actually replaced by ethanol, and the refined structure is a trinuclear SBU with MIL-53 (2D triangular) topology. The experiment therefore validates a MOF that is similar to, and inspired by, the model prediction, rather than the predicted structure itself. This distinction matters for a paper whose central claim includes a practical pathway to synthesizable MOFs, and it should be reflected in the abstract. In addition, the statement that the simulation 'fits well' with the PXRD data could be strengthened by reporting the refinement residuals (for example, Rwp and Rp).","section":"Abstract and Section 2d"},{"comment":"The edge-angle cutoff is stated inconsistently. Section 2a says edges with pairwise connecting-point angles of at least 120 degrees were used, while Sections 2b and 4b report a 140–180 degree constraint inherited from prior work and use that constraint to define edge validity. Because the distinction between curation criteria and validity criteria is central to the paper's quantitative claims, the applied cutoffs for training-data curation and for validity must be stated consistently; if two different cutoffs were in fact used, their relationship needs to be explained.","section":"Section 2a vs Sections 2b/4b"},{"comment":"The claims that generated MOFs 'faithfully represent' CoRE MOF and show 'highly similar distributions' of edge angles and node RMSDs (Figs. 2d and 3) are supported only by visual overlay, and the distribution comparison in Fig. 2d is computed on valid samples that are filtered by exactly the same 140-degree and 0.3 Å thresholds used to define validity. No quantitative distance measure (such as an RAC-space overlap, a KL divergence, or nearest-neighbor statistics) is reported, and no validity or novelty comparison against MOFDiff, MOFFlow, or another MOF generative baseline is provided. Without such comparisons, the abstract's claim that the model 'readily samples MOFs with unit cells containing 1000 atoms with great geometric validity, novelty, and diversity mirroring experimental databases' cannot be benchmarked against existing methods.","section":"Section 2c / Section 2b"}],"minor_comments":[{"comment":"Fig. 2c shows a maximum of 904 atoms in the unit cell, while the abstract and Section 2b mention '1000 atoms' and 'approximately 1000 atoms'; the numbers should be aligned.","section":"Abstract / Fig. 2c"},{"comment":"With 20% novel edges and 25% novel nodes among valid MOFs, a MOF containing at least one novel building block should occur more often than 25% under independence; the reported 27% overall novelty implies strong co-occurrence of novel edges and novel nodes (or a difference in the base sets over which the rates are computed), and this should be explained or clarified.","section":"Section 2b"},{"comment":"There are several typos in Section 2d, including 'tow carboxylates', 'An hexagonal node', and 'The N2 Adsorption-desorption Analysis was also experiment to obtained the BET surface areas', which should all be corrected.","section":"Section 2d"},{"comment":"The simplified training objective L_simple is written without the expectation over the data and noise distributions; the standard form with the expectation should be restored.","section":"Section 4a"},{"comment":"The conditionally generated edges in Fig. 2e are described qualitatively as 'reasonable'; a quantitative validity check for the inpainting products would strengthen the conditional-generation claim.","section":"Fig. 2e"},{"comment":"The edge-angle comparison between generated and CoRE MOF samples should either include invalid generated MOFs or explicitly note the truncation at the validity cutoff in the caption, since the present plot compares filtered generated samples against an unfiltered experimental distribution.","section":"Fig. 2d"}],"recommendation":"major_revision","confidential_remarks":"The methodological novelty relative to the authors' own OA-ReactDiff work lies mainly in the MOF building-block representation and its application; this is a legitimate contribution but should be positioned accordingly. The abstract overstates the experimental confirmation relative to what Section 2d actually reports; this needs reconciliation in revision. The 120-degree/140-degree cutoff inconsistency should be checked, and the authors should confirm that the deposited GitHub and Zenodo artifacts contain the BBA MOF Diffusion training and sampling pipeline rather than only the OA-ReactDiff code, since code availability is claimed as a strength. The paper's explicit discussion of its limitations is to its credit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is treating MOF generation as separate diffusion over nodes, edges, and topological nets, then assembling with PORMAKE. That lets the model propose novel building blocks instead of only recombining known ones, and it dodges the ~200-atom whole-cell limit—generated unit cells go up to ~900 atoms. The conditional inpainting demo with the paddle-wheel node is a nice practical touch, and the attempted synthesis, even with its blemishes, is more than most generative-model papers offer.\n\nWhere the paper is softer is the validity metric. Edge validity is defined by the same 140–180° angle cutoff used to curate the training set, and node validity by the same 0.3 Å RMSD threshold against idealized net positions. The paper itself admits that sulfonyldibenzene edges fall below 140° and still form valid MOFs. So the 52% overall validity is partly a restatement of the curation filters, not an independent measure of synthesizability. The abstract's \"great geometric validity\" is therefore doing more work than the metrics support. That said, the authors are not hiding this—Section 2b spells out the angle issue explicitly, and the supplement separates angle failures from structural-sanity failures. The fix is straightforward: report sanity rates separately, soften the abstract, and perhaps use a validity definition that does not recapitulate the training filters.\n\nThe synthesis section is honest but also dampens the claim. The model predicted a structure with two TDC ligands coordinating to Zn; the synthesized material has those replaced by ethanol. The authors disclose this and explain it via solvent loss during curation. That is fair, but it means the phrase \"structural fidelity\" in the abstract overstates what was actually confirmed. The PXRD refinement confirms the synthesized structure, not the predicted one. This is a moderate issue, not a fatal one.\n\nTwo technical gaps: no quantitative comparison to MOFDiff or MOFFlow, and the code link points to the OA-ReactDiff repository rather than a BBA-specific implementation. The data on Zenodo is good, but reviewers will want the actual model code. The paper also trains on all data without a test split—acceptable for a generative model whose output is judged by diversity and validity, but worth stating more clearly.\n\nThe RAC distribution comparison to CoRE and ToBaCCo is the most convincing part of the evaluation, and the scaling argument is solid. This is a serious, honest piece of work with a real methodological contribution. It deserves peer review, with the expectation that the authors add baselines, release the right code, and recalibrate the validity claim.","headline":"A useful building-block-aware diffusion framework for MOFs, with a real synthesis step, but the headline validity numbers inherit the curation cutoffs and the code link underdelivers.","tokens_in":19296,"tokens_out":1556,"would_cite":true,"duration_ms":15858,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By generating MOFs block by block, a diffusion model reaches 1,000-atom unit cells and proposes building blocks never seen in its training data.","keywords":["metal-organic frameworks","generative diffusion models","building-block representation","SE(3) equivariance","equivariant graph neural networks","topological nets","de novo materials design","CoRE MOF"],"falsifier":"Re-validate the generated MOFs using a lower edge-angle cutoff (e.g., 120 degrees) and a relaxed node RMSD (e.g., 0.5 Å), and count how many of the 52% valid structures remain valid; if the fraction drops sharply, the validity claim is an artifact of the curation thresholds. A second decisive test is to synthesize another high-scoring generated MOF with a novel node or edge and check whether the experimentally obtained structure matches the predicted framework, since the one successful synthesis involved a model structure whose solvent coordination differed from the experiment.","tokens_in":18154,"feed_emoji":"🧪","tokens_out":8394,"duration_ms":75792,"temperature":0.7,"pith_summary":"The paper tries to establish that a generative model can design new metal–organic frameworks (MOFs) by learning the three-dimensional shapes of their building blocks — inorganic nodes, organic edges, and topological nets — rather than by recycling known blocks or generating entire unit cells at once. If true, de novo MOF design can produce crystals with unit cells of roughly 1,000 atoms, a size all-atom diffusion has not previously reached, and can propose metal clusters and linkers absent from the training database. The authors report 52% geometric validity among assembled MOFs, 27% novelty, and diversity that overlaps experimental databases, and they synthesize one predicted MOF, confirming its overall framework by powder X-ray diffraction, thermogravimetric analysis, and nitrogen sorption.","feed_headline":"MOF generator invents new building blocks, scales to 1,000 atoms","feed_subtitle":"Building-block-aware diffusion samples novel nodes and linkers; one predicted MOF was synthesized and verified.","key_machinery":"The central machinery is the building-block representation itself, coupled with an object-aware $\\mathrm{SE}(3)$-equivariant diffusion framework. A MOF is disassembled into an inorganic node, an organic edge, and a topological net; the diffusion model denoises the all-atom Cartesian coordinates of each block while keeping atom types fixed, and the denoising network is a LEFTNet equivariant graph neural network adapted from the authors' earlier object-aware reaction diffusion model. Assembly is performed by PORMAKE, and conditional design is implemented by inpainting, which keeps known nodes or nets fixed while denoising only the target component. The reduction in graph size is what lets the model scale to unit cells of roughly 1,000 atoms.","core_discovery":"BBA MOF Diffusion is an $\\mathrm{SE}(3)$-equivariant denoising diffusion model that represents a MOF as a joint distribution over three objects: one inorganic node, one organic edge, and a topological net. Trained on the CoRE MOF 2019 database of experimentally synthesized MOFs, the model samples new all-atom nodes and edges and assembles them, via PORMAKE, onto one of four simple nets (dia, nbo, sra, pcu). The central claim is that this native building-block representation produces unprecedented metal nodes and organic edges, expanding accessible chemical space by orders of magnitude, while still yielding assembled MOFs with good geometric validity: 52% of sampled MOFs are valid, 27% are novel, 13% are both novel and unique, and unit cells range from 37 to 904 atoms. The synthesized [Zn(1,4-TDC)(EtOH)2] MOF demonstrates that at least one generated structure corresponds to a real material, with the experimental structure differing from the model's prediction in the coordination of ethanol solvent molecules.","pith_inferences":["A direct stress test of the validity claim would be to retrain or re-sample with the edge-angle cutoff lowered from 140 to about 120 degrees and re-measure validity; if the model's building-block geometries are genuinely sound, validity should degrade only mildly.","The model currently generates only the four most common nets and does not invent new topologies; treating the net itself as a diffused object, or sampling nets from a generative model of the long tail, would be the natural next step and is not tested in this paper.","Because the diffusion model keeps atom types fixed and only denoises coordinates, the chemistry of generated building blocks is bounded by the atomic compositions the user supplies; coupling this model with a composition generator would close the loop to fully automatic design.","The same object-aware, building-block treatment could transfer to other modular molecular systems where one component's identity and another's coordinates interact without direct 3D contact, such as co-crystals or protein–ligand complexes, although the paper does not demonstrate those cases."],"forward_implications":["Generative MOF design no longer needs to recycle known building blocks: the model can propose previously unseen metal nodes and organic edges, shifting the bottleneck from structure generation to synthesis planning.","Unit cells containing up to roughly 1000 atoms can be sampled, so all-atom generation can cover realistic MOF crystals rather than only small model systems.","Because the joint distribution is learned from experimentally synthesized MOFs, generated candidates resemble the diversity of CoRE MOF while adding novelty, making them plausible targets for experimental follow-up.","Conditional generation by inpainting lets chemists fix a known node and net and request linkers with a specified chemical composition, giving a practical workflow for linker-focused design.","At least one model-predicted MOF, [Zn(1,4-TDC)(EtOH)2], was synthesized and characterized, demonstrating that the pipeline can output a real, crystallographically confirmed material."],"supporting_citations":[{"why":"Supplies the CoRE MOF 2019 training database of experimentally synthesized MOFs.","marker":"[28]"},{"why":"Provides the automated deconstruction algorithm and the 0.3 Å RMSD and 140-degree cutoffs used to curate nodes and edges.","marker":"[34]"},{"why":"The object-aware SE(3)-equivariant diffusion framework that BBA MOF Diffusion adapts to MOF building blocks.","marker":"[62]"},{"why":"LEFTNet, the equivariant graph neural network used as the denoising scoring network.","marker":"[74]"},{"why":"PORMAKE, the assembly tool that reconstructs 3D MOF unit cells from generated nodes, edges, and nets.","marker":"[70]"},{"why":"MOFDiff, the coarse-grained diffusion baseline that recycles known building blocks and is contrasted with the de novo building-block generation here.","marker":"[64]"},{"why":"MOFFlow, a flow-matching baseline that also assembles existing building blocks, serving as the comparison for novelty.","marker":"[65]"},{"why":"RePaint inpainting, the technique used for conditional generation with fixed nodes or nets.","marker":"[82]"},{"why":"ToBaCCo, the topologically diverse hypothetical MOF database used as the comparator in the diversity analysis.","marker":"[33]"}],"fun_headline_variants":["AI invents novel MOF building blocks, scales to 1,000 atoms","Diffusion model designs MOFs with unprecedented nodes and linkers","MOF diffusion model samples 1,000-atom crystals, one synthesized","Building-block diffusion generates novel MOFs, verified by synthesis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the curation cutoffs used to build the training set — a 0.3 Å root-mean-square deviation for node–net compatibility and a 140-degree minimum angle between an edge's connecting points — define which building blocks are genuinely valid, so that the reported 52% validity reflects synthesizability rather than the curation rule itself; the paper concedes that some valid MOFs in CoRE have edge angles below 140 degrees.","fun_headline_variants_meta":{"raw":{"variants":["AI invents novel MOF building blocks, scales to 1,000 atoms","Diffusion model designs MOFs with unprecedented nodes and linkers","MOF diffusion model samples 1,000-atom crystals, one synthesized","Building-block diffusion generates novel MOFs, verified by synthesis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000593,"raw_usage":{"total_tokens":2797,"prompt_tokens":979,"completion_tokens":1818,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":1741}},"tokens_in":595,"tokens_out":1818,"duration_ms":11181,"temperature":1.0,"reasoning_tokens":1741,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:51:51.743389+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-validate the generated MOFs using a lower edge-angle cutoff (e.g., 120 degrees) and a relaxed node RMSD (e.g., 0.5 Å), and count how many of the 52% valid structures remain valid; if the fraction drops sharply, the validity claim is an artifact of the curation thresholds. A second decisive test is to synthesize another high-scoring generated MOF with a novel node or edge and check whether the experimentally obtained structure matches the predicted framework, since the one successful synthesis involved a model structure whose solvent coordination differed from the experiment.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the CoRE MOF 2019 training database of experimentally synthesized MOFs."},{"cited_title":", author Yue, S","cited_arxiv_id":null,"evidence_quote":"Provides the automated deconstruction algorithm and the 0.3 Å RMSD and 140-degree cutoffs used to curate nodes and edges."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ToBaCCo, the topologically diverse hypothetical MOF database used as the comparator in the diversity analysis."}],"review_version":1}