{"id":"6d64a567-e68f-4275-a2bf-a5708263b324","arxiv_id":"2504.12627","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"DPOSE uncertainty estimates on SchNet are higher for out-of-domain chemistry such as unseen elements, but are less reliable for structurally similar out-of-domain configurations.","lead":"This paper integrates a lightweight uncertainty method called DPOSE into the SchNet machine learning potential and tests whether the predicted uncertainty rises for molecules and materials outside the training data. It finds that uncertainty grows for unseen elements and larger molecules, but not always for structural changes like bulk versus amorphous gold.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never ties head variance to per-sample error on OOD data; without a calibration check, the reported domain-averaged variance gaps could be a target-scale artifact rather than evidence that DPOSE produces usable uncertainty.","rationale":"The paper's stated purpose is to detect OOD inputs via uncertainty. That is a claim about the relationship between variance and prediction error, not merely about average variance differences between two curated groups. The QM9 and gold composition results are suggestive, and the paper honestly acknowledges the gold bulk/amorphous limitation (Sections 3.3.1 and 4). The OC20 and QM9 numbers, however, are not accompanied by calibration curves, error correlations, or per-atom normalization. Since total energy and its head-variance both scale with system size, the monotonic size trend in Table 2 could be explained without any increase in per-atom uncertainty. This is a concrete, testable confound, and it is exactly the kind of check that would tell whether the conditional verdict should become acceptance or rejection. I therefore agree with the reader's conditional verdict and with the weakest-assumption diagnosis: the missing validation is that head variance is a reliable error proxy on shifted data. No verdict adjustment is needed beyond the existing conditionality.","tokens_in":10149,"tokens_out":6188,"duration_ms":69041,"concrete_test":"On the QM9 out-of-domain set, collect per-molecule 64-head variance and per-molecule absolute prediction error for all substituted molecules (Si/S/P/Cl) and C1–C18 alkanes. Compute Spearman rank correlation between variance and squared error, first on raw total energies and again after dividing variance and error by N_atoms. Also compute a binned calibration curve (mean squared error vs variance decile). If the rank correlation is weak (rho < 0.5) or the C1–C18 variance trend disappears under per-atom normalization, the reported OOD separation is a scale artifact, and the central claim fails. If per-atom variance still increases with size and orders errors within the OOD set, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that DPOSE/SchNet 'successfully distinguishes in-domain from out-of-domain data'—requires the 64-head variance from the NLL loss (Eq. 1) to be a usable proxy for predictive error on shifted inputs. The paper only shows that hand-picked domain groups have different average variances (CF4 0.09 vs CCl4 62; Au 0.00022 vs Ag 0.613 eV/atom). It never checks the operative property: whether on OOD inputs variance orders samples by actual error, or whether the large gaps are partly artefactual. Tables 1–2 report total-energy variances, and energy is extensive: dodecane (0.62) and octadecane (0.66) exceed nonane (0.11) by more than a factor of 5, yet the trend is non-monotonic and no per-atom normalization is given. The same scale confound could affect the OC20 variance comparisons if slab sizes differ. Because the conclusion frames these as 'reliable uncertainty estimates' for active learning, the missing error-vs-variance calibration is the load-bearing gap: average differences can be large even if head variance is uninformative for individual OOD samples.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper integrates Direct Propagation of Shallow Ensembles (DPOSE) into the SchNet graph neural network by replacing the final layer with 64 output heads and training with the negative log-likelihood loss of Eq. (1). The resulting variance of the heads is used as an uncertainty estimate. The model is evaluated on three datasets: QM9 (equilibrium vs. non-equilibrium geometries, known vs. unknown elements, small vs. large molecules), OC20 (intermetallic vs. non-metal slabs, volume compression/expansion), and a Gold MD dataset (bulk vs. amorphous, Au vs. Ag). The reported results show large variance gaps for compositional shifts (e.g., CF4 0.09 vs. CCl4 62; Au 0.00022 vs. Ag 0.613 eV/atom), a gradual variance increase with molecular size, higher variance for non-metal slabs, and essentially no variance difference between bulk and amorphous gold. The authors conclude that DPOSE distinguishes in-domain from out-of-domain data when the domains are far apart, while acknowledging difficulty for structurally similar systems.","tokens_in":10419,"tokens_out":3979,"duration_ms":42208,"significance":"If head variance is a reliable proxy for predictive error, this would be a useful lightweight UQ method for GNN potentials, particularly for active learning. The study has strengths: it tests three diverse datasets, examines both compositional and structural shifts, and is transparent about a negative result (bulk vs. amorphous). However, the central claim of 'reliable uncertainty estimates' is not quantitatively validated, and the reported variance differences are aggregate and potentially scale-confounded. The contribution is an exploratory empirical study; its value depends on a calibration check and comparison with existing baselines.","major_comments":[{"comment":"The central claim that DPOSE provides reliable uncertainty estimates requires demonstrating that the head variance from Eq. (1) is predictive of predictive error on individual out-of-domain samples. All reported evidence consists of aggregate variance differences between hand-picked groups (Tables 1–2, Figures 3–9); there is no calibration metric such as a reliability diagram, Spearman rank correlation between variance and absolute error, or an error-vs-variance plot for OOD samples. Without such a check, the large average gaps could be a feature-scale artifact rather than evidence of usable uncertainty.","section":"§2.2, §3 (general)"},{"comment":"The claim that uncertainty increases with molecular size is confounded by the extensivity of total energy. The QM9 variance values are reported without units or per-atom normalization, and the trend is non-monotonic (dodecane 0.62, tridecane 0.82, tetradecane 0.61). Since larger molecules have larger total energies, a naive variance of total energy may scale with system size even when per-atom uncertainty is constant. The authors should report normalized per-atom variances or otherwise show that the size trend is not a scale effect.","section":"§3.1.3, Table 2"},{"comment":"The bulk-vs-amorphous result contradicts the stated expectation of lower uncertainty for in-domain bulk systems; the text acknowledges 'the model seems to be confident for out-of-domain systems too' and later that DPOSE 'struggled to distinguish between in-domain and out-of-domain configurations.' This is a central negative result for structural shifts. The abstract and conclusion currently over-generalize by saying DPOSE 'often demonstrates' or 'successfully distinguishes' without clearly carving out that the supported success is limited to compositional shifts and large domain distances. The paper should explicitly state in the abstract that structural OOD detection failed for Au bulk vs. amorphous.","section":"§3.3.1, Figure 6(c), Conclusion"},{"comment":"The paper motivates DPOSE as a computationally efficient alternative to deep ensembles but never compares against any baseline UQ method, such as deep ensembles, MC dropout, or latent-distance uncertainty. Without a baseline, the strong wording about 'reliable' and 'better-calibrated' uncertainty is unsupported. At minimum, the authors should include one standard baseline with a calibration metric on the same OOD tasks.","section":"§2, §3 (no baseline comparison)"}],"minor_comments":[{"comment":"Equation (1) is presented without derivation or an explicit connection to the original DPOSE paper [19]; please cite the source in the text and define all symbols (e.g., whether sigma is per-output or shared).","section":"§2, Eq. (1)"},{"comment":"The name 'SchNET' is inconsistent with the standard 'SchNet' used elsewhere; correct the typo.","section":"§2.1"},{"comment":"The QM9 variance values are reported without units and without stating whether they are per atom or total. Please specify units (e.g., eV^2, (kcal/mol)^2) and clarify normalization.","section":"Tables 1–2, §3.1"},{"comment":"The caption refers to a 'red line' while the text and figure show a red dotted line; make the reference consistent.","section":"Figure 5, caption"},{"comment":"The terms 'inter-metallic' and 'intermetallic' are used inconsistently (e.g., §3.2.1 'inter-metallic slabs' vs. §2.2.2 'intermetallic slab subset'); choose one spelling.","section":"§3.2, consistent terminology"},{"comment":"The text says 'the mean variance was slightly higher for bulk systems' while the caption says 'Variance estimates showing slightly higher uncertainty for bulk'; clarify whether the box plot displays means, medians, or distributions.","section":"Figure 6(c), §3.3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for the journal but is more of an empirical exploration than a rigorous UQ evaluation. The authors are honest about the negative result for bulk vs. amorphous, which is commendable. The main missing ingredient is any quantitative calibration or error-correlation test; without it, the central claim cannot be evaluated. I would suggest requesting code/data availability and a baseline comparison in the revision; this is not a citation-pattern concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is not a methods paper—DPOSE already exists (Kellner & Ceriotti). The contribution is the application to SchNet across three datasets, and one genuinely useful empirical observation: composition shifts (unseen elements) produce large variance gaps, while structural shifts (bulk vs amorphous gold) do not. The paper is candid about the latter failure, which adds credibility.\n\nWhat's good: the qualitative trends for QM9 are stark—CF4 variance 0.09 vs CCl4 62, with equal atom counts, so not a size artifact. Same for Au vs Ag (0.00022 vs 0.613 eV/atom). Those results are hard to dismiss. The energy-variance-error plots for gold also show a sensible monotone relationship. The writing is clear, and the conclusion is appropriately hedged.\n\nSoft spots, in decreasing order:\n\n1. No calibration or ranking metrics. The central claim is that variance flags OOD samples. The paper never checks whether variance orders individual samples by actual error, or whether average gaps are driven by a few outliers. The stress-test note is right: domain-averaged variance differences could be large even if per-sample variance is uninformative. This is the load-bearing gap.\n\n2. The QM9 variance values appear to be total-energy variances, not per-atom. Table 2 shows a jump from C9 (0.11) to C12 (0.62) that is not simply linear in atom count, but without per-atom normalization the size trend is confounded. They use per-atom units for Gold and OC20 but not QM9. Easy fix.\n\n3. No baselines. They motivate DPOSE as efficient vs deep ensembles, but never compare runtime, accuracy, or uncertainty quality against deep ensembles, MC dropout, or latent distance. The empirical observation stands on its own, but the contextual claims are unsubstantiated.\n\n4. Minor: the 80:20 split description conflicts with 'trained on the entire QM9 dataset'; some numbers are given without error bars (single molecule variances). These are small.\n\nOverall: I believe the paper reports real effects for composition shifts, and the negative result on structural shifts is honest. But it reads like a workshop paper or an exploratory report, not a complete study. With calibration metrics, per-atom normalization, and one baseline, it could be a solid contribution to the atomistic UQ literature.\n\nRecommendation: send it to peer review—a serious referee can push the authors to add the missing analysis. It is not a desk reject; the core evidence is worth taking seriously.","headline":"A light but honest empirical study of DPOSE on SchNet; the composition-shift results are convincing, but the lack of calibration and baselines keeps it a preliminary exploration.","tokens_in":10900,"tokens_out":2628,"would_cite":false,"duration_ms":27233,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Shallow ensembles of 64 SchNet heads can flag out-of-domain molecules and materials through prediction variance alone, when the domain shift is large.","keywords":["uncertainty quantification","shallow ensembles","graph neural networks","SchNet","out-of-domain detection","machine-learned potentials","negative log-likelihood loss"],"falsifier":"Take the QM9 substitution experiments and build a reliability diagram: bin molecules by DPOSE variance and plot the mean absolute prediction error in each bin; if the highest-variance molecules such as CCl4 and SiCl4 do not also show the largest errors, or if variance and error rankings diverge across the full set, the variance-as-uncertainty interpretation would fail even though the plotted variance trends remain.","tokens_in":10009,"feed_emoji":"⚛️","tokens_out":9111,"duration_ms":77856,"temperature":0.7,"pith_summary":"This paper tries to establish that a lightweight uncertainty-quantification method, Direct Propagation of Shallow Ensembles (DPOSE), can make graph neural network potentials tell users when they are extrapolating. The authors attach 64 shared-weight output heads to SchNet, train them with a negative log-likelihood loss, and use the spread of the heads' energy predictions as the uncertainty estimate. They report that the spread stays low for in-domain inputs and rises sharply for unseen elements, larger molecules, and chemically different slabs, which matters because materials discovery depends on knowing when a prediction is unreliable. The paper also reports a clear limit: the same method did not separate amorphous gold from bulk gold, so it flags compositional novelty more readily than same-element structural novelty.","feed_headline":"64 heads, one loss: variance flags unseen elements and molecules","feed_subtitle":"A shallow ensemble flags when a materials model is extrapolating, so searches can skip bad predictions.","key_machinery":"The central object is DPOSE, a shallow-ensemble method in which the model's last layer is replaced by 64 parallel output heads that share all earlier weights; the mean of the heads is the prediction and their variance is the uncertainty. The heads are trained jointly with the negative log-likelihood loss $\\text{NLL}(\\Delta y,\\sigma) = \\frac{1}{2}\\left(\\frac{\\Delta y^2}{\\sigma^2} + \\ln(2\\pi\\sigma^2)\\right)$, which couples each head's squared error to the ensemble variance. This mechanism lets one SchNet pass produce both an energy estimate and an uncertainty estimate, with a computational cost the paper compares to adding an extra hidden layer.","core_discovery":"On its own terms, the paper claims that DPOSE-equipped SchNet produces useful out-of-domain signals without the cost of deep ensembles. The variance across 64 output heads increases from 0.09 for CF4 to 62 for CCl4, from 0.09 for CH4 to 2.7 for SiH4, and from 0.00022 eV/atom for gold to 0.613 eV/atom for silver, while alkane uncertainties grow from 0.09 for methane to 0.66 for octadecane as chain length leaves the training range. In OC20, non-metal slabs show median variance near 0.03 while intermetallic slabs stay below 0.01. The authors conclude that DPOSE distinguishes in-domain from out-of-domain data when the two are far apart, but that structurally similar systems made of the same element, such as amorphous versus bulk gold, are not separated by variance.","pith_inferences":["Inference: the observed contrast suggests DPOSE variance encodes compositional novelty more strongly than geometric novelty; a testable extension is to correlate variance with element-embedding distance rather than with structural descriptors.","Inference: since the 64 heads share all weights except the last layer, the ensemble diversity is limited to the final projection, so the variance may reflect last-layer sensitivity more than full epistemic uncertainty; comparing head counts (e.g., 8, 16, 64) would show how much of the signal comes from width alone.","Inference: the paper does not calibrate variance against prediction error; a natural extension is to check whether the reported variance ordering matches the actual error ordering on the same out-of-domain molecules, which would turn a qualitative trend into a usable uncertainty score.","Inference: for active learning, a DPOSE-variance acquisition function would likely prioritize novel elements and long molecules first, while missing amorphous phases of known elements; pairing variance with a latent-distance or energy-based score could cover that blind spot."],"forward_implications":["A single SchNet model with 64 heads can emit an uncertainty flag for every energy prediction, so out-of-domain detection does not require training several full models.","Compositional shifts, such as substituting an unseen element into a known molecule, produce variance jumps large enough to separate in-domain from out-of-domain molecules in QM9 and in gold-versus-silver comparisons.","Uncertainty tracks physical extrapolation distance, rising as carbon-carbon bonds are stretched or compressed and as alkane or alcohol chains grow beyond the training lengths.","The method's failure to separate amorphous from bulk gold implies that same-element structural diversity is not reliably flagged, so the practical alarm is for chemistry novelty, not morphology novelty.","Because the added cost is comparable to one extra hidden layer, the approach is light enough to be coupled with active learning loops that acquire new training data where variance is high."],"supporting_citations":[{"why":"Introduces DPOSE, the shallow-ensemble with weight sharing and NLL loss that the paper adapts to SchNet.","marker":"[19]"},{"why":"Supplies the SchNet graph neural network architecture whose final layer is split into 64 heads.","marker":"[20]"},{"why":"Provides the QM9 dataset of small organic molecules used for the equilibrium, element-substitution, and molecule-size tests.","marker":"[22]"},{"why":"Complements [22] as the source of the QM9 molecular structures and properties.","marker":"[23]"},{"why":"Provides the OC20 dataset of slabs and adsorbates used to compare intermetallic and non-metal systems.","marker":"[21]"},{"why":"Provides the gold molecular dynamics dataset used for the bulk-versus-amorphous and gold-versus-silver tests.","marker":"[24]"}],"fun_headline_variants":["Variance from 64 heads flags unseen atoms and molecules","Shallow ensemble variance exposes extrapolation in GNNs","DPOSE: lightweight variance for out-of-domain detection","Variance jumps reveal when materials GNNs leave training set"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the variance among the 64 NLL-trained heads is a trustworthy stand-in for predictive error on out-of-domain inputs, which enters when the paper reads higher variance as lower confidence and is never checked against calibration.","fun_headline_variants_meta":{"raw":{"variants":["Variance from 64 heads flags unseen atoms and molecules","Shallow ensemble variance exposes extrapolation in GNNs","DPOSE: lightweight variance for out-of-domain detection","Variance jumps reveal when materials GNNs leave training set"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001173,"raw_usage":{"total_tokens":4838,"prompt_tokens":918,"completion_tokens":3920,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":3853}},"tokens_in":534,"tokens_out":3920,"duration_ms":29561,"temperature":1.0,"reasoning_tokens":3853,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:26:28.933770+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the QM9 substitution experiments and build a reliability diagram: bin molecules by DPOSE variance and plot the mean absolute prediction error in each bin; if the highest-variance molecules such as CCl4 and SiCl4 do not also show the largest errors, or if variance and error rankings diverge across the full set, the variance-as-uncertainty interpretation would fail even though the plotted variance trends remain.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the QM9 dataset of small organic molecules used for the equilibrium, element-substitution, and molecule-size tests."},{"cited_title":"wiley.com/doi/abs/10.1002/qua.25115","cited_arxiv_id":null,"evidence_quote":"Provides the gold molecular dynamics dataset used for the bulk-versus-amorphous and gold-versus-silver tests."}],"review_version":1}