{"id":"438cc04a-d351-43f7-bed4-de523406d6cf","arxiv_id":"2601.13125","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Underestimation of PbTiO3's Curie temperature is dominated by the exchange-correlation functional; short-range machine-learning potentials appear closer to experiment only through cancellation of errors.","lead":"Using large-scale simulations, this paper shows that machine-learning force fields reproduce ab initio results for PbTiO3, so the underestimation of the Curie temperature comes from the density-functional approximation itself, not from the machine learning. The findings clarify how to build reliable interatomic potentials for ferroelectric materials.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"qNEP's 600 K 'true PBEsol limit' is not thermodynamically validated: force RMSE alone does not establish that the latent-charge model captures the free-energy surface at 12^3 supercells.","rationale":"I read the paper in good faith. Its strongest contribution is the demonstration that a well-trained DP model reproduces the AIMD Tc and polarization behavior at 320 atoms, thereby shifting the blame for Tc underestimation from MLFF fitting to the underlying PBEsol functional. That part is convincing and is independently supported by direct AIMD. The more ambitious statement that the true PBEsol thermodynamic limit is exactly 600 K depends on qNEP's ability to represent long-range electrostatics correctly at 12x12x12. The evidence for qNEP is force RMSE, which is necessary but not sufficient for thermodynamic accuracy. The latent-charge scheme introduces a model assumption that could systematically shift the free-energy landscape. The reader's weakest assumption already identified the lack of direct AIMD verification for qNEP at large sizes; my concern sharpens this by noting that force RMSE does not validate the thermodynamic quantity being predicted. The internal discrepancy between 450 and 500 K for the same small-cell Tc is a secondary but real inconsistency that the authors should address. None of this overturns the paper's main qualitative conclusion, so the existing CONDITIONAL verdict remains appropriate rather than a more severe change.","tokens_in":10759,"tokens_out":4439,"duration_ms":47814,"concrete_test":"Run qNEP NPT MD for the same 320-atom 4x4x4 supercell using the same protocol as the AIMD simulations. If qNEP reproduces AIMD's ~500 K Tc and the Pz(T), c/a(T) curves at this size, then the long-range model is thermodynamically validated at the only size with AIMD ground truth. Then run qNEP at 6x6x6 and 8x8x8 supercells to confirm a monotonic finite-size trend toward 600 K. If qNEP at 320 atoms disagrees with AIMD, the 600 K conclusion should be treated as unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's qualitative central claim — that MLFF fitting error is not the primary cause of Tc underestimation — is well supported by the direct AIMD/DP agreement at 320 atoms. The load-bearing weak point is the quantitative claim that 600 K is the 'true thermodynamic limit of PBEsol.' That claim rests entirely on qNEP simulations at 12x12x12. The only validation provided for qNEP (Fig. 6a) is force RMSE against DFT on supercells up to 5000 atoms. Force RMSE is a local metric; Tc is controlled by the free-energy difference between ferroelectric and paraelectric phases, and a model can have low force errors yet incorrect anharmonic free-energy curvature. qNEP's latent charges are fit without any direct DFT charge or electrostatic reference, so its lower Tc relative to local NEP could be a modeling artifact rather than the removal of short-range truncation error. Moreover, no direct AIMD comparison at 320 atoms is shown for qNEP, and no AIMD exists at larger sizes, so the 'true thermodynamic limit' statement is an extrapolation. An additional internal inconsistency — the small-cell Tc is reported as 500 K in Sec. III.A–C but appears as 450 K in Sec. III.E — further weakens confidence in the finite-size convergence narrative. The qualitative conclusion likely survives, but the specific 600 K limit is conditional and not yet established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses NPT AIMD (320-atom supercell, PBEsol) and MLFF benchmark simulations to explain why computed Curie temperatures of PbTiO3 fall below the experimental 760 K. AIMD and a previously trained DP model both give Tc ≈ 500 K at 320 atoms, which the authors interpret as showing that the MLFF fitting is not the cause of the underestimation. DP simulations for larger supercells show Tc increasing to about 650 K. To probe the role of long-range electrostatics, the authors train NEP and qNEP models and find, in 12×12×12 supercells, that short-range NEP gives Tc ≈ 650 K while qNEP with explicit latent-charge electrostatics gives Tc ≈ 600 K. They conclude that 600 K is the true thermodynamic limit of PBEsol for PbTiO3, and that the better agreement of short-range models arises from a fortuitous cancellation of errors.","tokens_in":11121,"tokens_out":2751,"duration_ms":34075,"significance":"If the conclusions hold, the paper makes a useful contribution by separating three sources of error in first-principles Tc predictions: MLFF fitting error, finite-size effects, and exchange-correlation functional error. The qualitative attribution of the underestimation to PBEsol is well supported by the direct agreement between AIMD and an independently trained DP model. The work also has practical value: the AIMD trajectories and trained models are made publicly available, and the paper offers a clear warning that agreement with experiment can result from error cancellation. However, the quantitative claim that 600 K is the true PBEsol thermodynamic limit rests on qNEP simulations whose validation is currently incomplete, so that part of the paper is conditional rather than established.","major_comments":[{"comment":"There is an internal inconsistency in the reported small-supercell Tc. Sections III.A–C and Figs. 1–3 report Tc ≈ 500 K from AIMD and DP at 320 atoms, but §III.E states that the small supercell gives 450 K when comparing with the 650 K large-cell result. The magnitude of the finite-size shift therefore changes depending on which value is used. This inconsistency is load-bearing for the finite-size convergence narrative and must be corrected and reconciled.","section":"§III.E (Fig. 5)"},{"comment":"The claim that 600 K is the true thermodynamic limit of PBEsol is not yet supported by the evidence presented. The only validation of qNEP shown in Fig. 6a is force RMSE against DFT on supercells of increasing size. Force RMSE is a local, zero-temperature metric; the Curie temperature is controlled by the free-energy difference between ferroelectric and paraelectric phases, including anharmonic contributions that a low force RMSE does not guarantee. The latent-charge model is not validated against any DFT charge or electrostatic reference, and no direct comparison of qNEP to AIMD at the 320-atom size is provided. Moreover, only a single 12×12×12 supercell is used for qNEP; there is no size-convergence series for Tc and no statistical uncertainty estimate. I request additional evidence: qNEP versus AIMD at 320 atoms for c/a and polarization distributions, a Tc convergence series (e.g., 8×","section":"§III.F (Fig. 6)"},{"comment":"The validation of DP_AIMD is circular: the model is trained exclusively on AIMD trajectory configurations and then compared to those same AIMD results. This demonstrates that the model can fit and reproduce its training data, but it does not demonstrate transferability or independent accuracy. This does not undermine the central argument, because the main DP model is trained on the DPGEN dataset and is compared to separate AIMD configurations, but the text should explicitly distinguish these two cases so that DP_AIMD is not presented as independent validation.","section":"§III.D (Fig. 4c)"}],"minor_comments":[{"comment":"The NEP and qNEP models are not described with the same level of detail as the DP model. Hyperparameters, descriptor settings, training-set sizes, and simulation protocols (thermostat, barostat, equilibration time, production length) for the NEP/qNEP MD runs in Fig. 6b should be given, either in the text or in the repository.","section":"§II.B"},{"comment":"The curves for different supercell sizes are not labeled directly; a legend identifying 4×4×4, 6×6×6, 8×8×8, and 10×10×10 is needed. Error bars on c/a and on the inferred Tc values would also help assess the convergence claim.","section":"Fig. 5a"},{"comment":"The figure reports Tc for NEP and qNEP but no statistical uncertainty or simulation-length information is provided. At minimum, the number of independent runs and the length of the production trajectories should be stated, since Tc is extracted from curves without error bars.","section":"Fig. 6b"},{"comment":"The sentence 'To the best of our knowledge, this represents the largest AIMD study of PbTiO3 to date' is a strong claim. It can be retained if a literature search has been performed, but a citation to the prior largest study would make the claim verifiable.","section":"§III.A"},{"comment":"The word 'conclusively' in the conclusion is stronger than the evidence supports, particularly given the incomplete qNEP validation. Softer wording would be more appropriate.","section":"Abstract/§IV"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a question of current interest and the central qualitative conclusion — that the PBEsol functional, not MLFF fitting error, drives the low Tc — is likely correct and well supported by the AIMD/DP agreement at 320 atoms. The main weakness is the overreach in the claim that 600 K is the true thermodynamic limit of PBEsol: this rests on a single qNEP simulation size and a validation metric (force RMSE) that is not sufficient for free-energy predictions. I recommend major revision rather than rejection because the missing tests are within the scope of a revision and the qualitative message will probably survive. The 450 K versus 500 K inconsistency must also be resolved; as written, it undermines confidence in the finite-size analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper disentangles two error sources in MLFF-based Tc prediction for PbTiO3: fitting error and functional error. The direct AIMD work is the real contribution—320-atom NPT simulations at 50 ps plus, which is substantial for CP2K—and the AIMD/DP agreement at Tc ≈ 500 K is convincing evidence that the MLFF is not the bottleneck. The catch that short-range models look better only because finite-size suppression cancels local-model stiffening is a genuinely useful insight, and the size-dependent force RMSE data for NEP vs qNEP support it. I also appreciate the DPGEN-vs-AIMD dataset comparison, though the DP_AIMD model trained on the same trajectory it is validated against is a weaker test than the paper implies; it shows self-consistency, not transferability. That circularity does not hurt the main conclusion, because the AIMD itself already gives the low Tc.\n\nThe soft spots are in the quantitative claim that 600 K is the true thermodynamic limit of PBEsol. That claim rests entirely on qNEP at 12×12×12, with validation limited to force RMSE against DFT on configurations up to 5000 atoms. Force RMSE is a local metric; it does not guarantee that the latent-charge model gets the anharmonic free-energy surface right, and qNEP’s charges are fit without direct DFT electrostatic reference. There is also no direct AIMD comparison at any size to confirm that qNEP reproduces the PES, and no statistical uncertainties on the Tc estimates. The internal inconsistency—small-cell Tc reported as 500 K in Sections III.A–C but 450 K in Section III.E—further undermines confidence in the finite-size narrative. These are fixable in revision: add a size-convergence series for qNEP, show at least one direct AIMD comparison for qNEP at 320 atoms, provide error bars, and clarify the 450/500 discrepancy.\n\nWho is this for? Anyone modeling phase transitions in ferroelectrics, and anyone building MLFFs for polar materials. The main qualitative result—functional error dominates over MLFF fitting error—is important and likely to stand. The specific 600 K limit should be treated as conditional until the qNEP evidence is strengthened.\n\nRecommendation: send it to peer review. A serious referee should push on the qNEP validation and the missing error analysis, but the paper deserves engagement.","headline":"Solid central result—Tc underestimation is a functional problem, not a fitting problem—but the specific 600 K PBEsol limit rests on unvalidated qNEP extrapolation and needs scrutiny.","tokens_in":11554,"tokens_out":1571,"would_cite":true,"duration_ms":18511,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["77.80.Bh"],"model":"deepseek-v4-flash","headline":"The paper shows that the underestimation of PbTiO3's Curie temperature in first-principles simulations comes from the intrinsic limits of the PBEsol exchange-correlation functional, not from machine-learning force-field fitting, and that wi","keywords":["Ferroelectricity","Curie temperature","PbTiO3","ab initio molecular dynamics","machine-learned force fields","long-range electrostatics","finite-size effects","exchange-correlation functional"],"falsifier":"Run constant-pressure ab initio molecular dynamics with PBEsol on an 8x8x8 or larger supercell (thousands of atoms) for sufficiently long trajectories near the expected transition: if the polarization order parameter disappears above about 650 K, the 600 K limit is wrong. Alternatively, perform the same 12x12x12 simulation with a hybrid functional; if Tc rises substantially beyond 600 K, the attribution to the exchange-correlation functional is supported, while if it stays near 600 K, the latent-charge model's extrapolation is suspect.","tokens_in":10663,"feed_emoji":"🌡️","tokens_out":4364,"duration_ms":42653,"temperature":0.7,"pith_summary":"The paper asks why theoretical Curie temperatures for PbTiO3 fall far below the experimental 760 K. By running large constant-pressure ab initio molecular dynamics simulations and benchmarking machine-learned force fields against them, it isolates the source of the error. The machine-learned potential reproduces the underlying density-functional calculation almost exactly, yielding the same transition temperature as direct simulation (about 500 K in a 320-atom cell). The gap with experiment therefore rests with the PBEsol functional itself. The paper further shows that a purely local machine-learned potential appears to raise the transition temperature to 650 K when the simulation cell is enlarged, but that rise is an artifact of neglected electrostatics; including long-range Coulomb interactions explicitly brings the converged prediction down to about 600 K, which the paper identifies as the true PBEsol thermodynamic limit.","feed_headline":"PBEsol caps PbTiO3's Curie temperature at 600 K","feed_subtitle":"Machine-learned potentials trace the 160 K gap with experiment to the functional itself — closing it needs new functionals.","key_machinery":"The argument hinges on comparing two classes of machine-learned potentials: a short-range model whose descriptors encode only local atomic environments, and a long-range variant that augments the short-range network with environment-dependent partial charges to compute explicit Coulomb interactions (a latent-charge model). The long-range model is the load-bearing instrument: it removes the artificial stiffening caused by electrostatic truncation and allows the paper to read off the functional's converged transition temperature in large supercells.","core_discovery":"The persistent underestimation of Tc in ferroelectric PbTiO3 originates primarily from intrinsic limitations of the PBEsol exchange-correlation functional, not from inaccuracies in the machine-learned force field. A deep-neural-network potential trained on diverse density-functional data reproduces the ab initio potential energy surface faithfully, giving the same Tc of about 500 K in a 4x4x4 supercell as direct AIMD. Increasing supercell size with short-range local models raises Tc to about 650 K, but this apparent improvement is a fortuitous cancellation of errors: local descriptors truncate electrostatics and artificially stiffen the lattice. When long-range electrostatics are explicitly","pith_inferences":["If the latent-charge model transfers as claimed, the same decomposition of errors (functional, model, finite-size) could be applied to other displacive ferroelectrics such as BaTiO3, where similar Tc underestimations have been reported.","The 600 K versus 760 K gap quantifies how much phase-transition physics is missing from semilocal DFT; a hybrid-functional AIMD in a comparably large cell would directly test whether the residual gap is indeed the functional's responsibility.","The result warns that matching experimental Tc with a local machine-learned potential is not evidence of accuracy; a model with larger test-set errors may appear more accurate for the transition temperature.","The observation that a long-range model achieves even lower errors for large supercells than for the 320-atom training cell suggests that finite-size artifacts in small cells are a separate challenge that explicit electrostatics resolves naturally."],"forward_implications":["Machine-learned potentials trained on diverse DFT data will faithfully reproduce the training functional's transition temperature, so fitting error is not the cause of the experimental gap in PbTiO3.","Short-range potentials systematically accumulate force errors as the supercell grows, while potentials with explicit long-range electrostatics remain accurate across sizes.","The 600 K value represents the PBEsol limit for PbTiO3, so closing the gap to 760 K requires improved exchange-correlation functionals beyond the generalized gradient approximation.","Small supercells (320 atoms) suppress Tc by roughly 100 K compared to the converged limit, so finite-size convergence must be checked in ferroelectric phase-transition simulations.","Accurate finite-temperature predictions need high-quality training data, large simulation cells, and explicit treatment of long-range interactions."],"fun_headline_variants":["Ferroelectric Tc gap traced to functional, not ML force field","PBEsol, not the ML model, caps PbTiO3's Tc","Why PbTiO3's predicted Tc falls short: functional fault","Short-range ML models mask PbTiO3 Tc error via cancellation","Long-range interactions essential for accurate PbTiO3 Tc"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The 600 K result assumes that the latent-charge long-range model, trained on 320-atom supercells, remains accurate in 12x12x12 supercells and truly represents the infinite-size PBEsol potential energy surface; if its long-range electrostatics drift at larger sizes, the 600 K limit is an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Ferroelectric Tc gap traced to functional, not ML force field","PBEsol, not the ML model, caps PbTiO3's Tc","Why PbTiO3's predicted Tc falls short: functional fault","Short-range ML models mask PbTiO3 Tc error via cancellation","Long-range interactions essential for accurate PbTiO3 Tc"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000207,"raw_usage":{"total_tokens":1242,"prompt_tokens":756,"completion_tokens":486,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":408}},"tokens_in":500,"tokens_out":486,"duration_ms":4982,"temperature":1.0,"reasoning_tokens":408,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T09:36:26.411121+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run constant-pressure ab initio molecular dynamics with PBEsol on an 8x8x8 or larger supercell (thousands of atoms) for sufficiently long trajectories near the expected transition: if the polarization order parameter disappears above about 650 K, the 600 K limit is wrong. Alternatively, perform the same 12x12x12 simulation with a hybrid functional; if Tc rises substantially beyond 600 K, the attribution to the exchange-correlation functional is supported, while if it stays near 600 K, the latent-charge model's extrapolation is suspect.","supporting_citations":[],"review_version":1}