{"id":"a20863f4-d05b-49b7-85b8-f4c420111208","arxiv_id":"2607.06645","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A hierarchical ALIGNN + ALIGNN-FF + DFT pipeline recovers experimental average voltages within 0.3 V on four commercial Li-ion cathodes and ranks large materials databases into surrogate-level shortlists.","lead":"BatteryMat is a three-tier pipeline that screens crystal databases for lithium-ion cathodes by predicting average voltage with ALIGNN, then refining survivors with a force field and automated DFT. It recovers experimental voltages of four commercial cathodes within 0.3 V and produces shortlists of prioritised candidates from JARVIS-DFT and Alexandria.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The primary screening signal is a distillation of ALIGNN-FF labels whose fidelity to DFT vacancy ordering and plateaus is only weakly checked outside five known cathodes.","rationale":"The reader correctly identifies the ALIGNN-FF fidelity assumption as the weakest link and already assigns CONDITIONAL with medium correctness risk. The DFT-tier experimental agreement on four stoichiometric cathodes is concrete and carefully caveated (NMC edge case, opposite-sign olivine residuals, finite-supercell plateaus). The methodological care on the Li reference (~1 V offset removed) and spacegroup-based functional selection is real. The shortlists are explicitly labelled surrogate-level leads, not confirmed cathodes. No stronger internal inconsistency appears: the 0.17 V number is not oversold. The remaining concern is exactly the one the reader flags—whether the force-field distillation remains a useful ranking signal outside the five known chemistries. That does not overturn the methodological contribution or the four-chemistry benchmark, so the verdict stays CONDITIONAL rather than REJECT; the concrete vacancy-order / step-voltage comparison on the actual shortlist heads is the check that would decide whether the shortlists can be treated as high-confidence leads. Agreement with the reader is therefore full on both the load-bearing soft spot and the recommended posture.","tokens_in":22160,"tokens_out":795,"duration_ms":8383,"concrete_test":"Take the top-5 JARVIS shortlist entries in Table 2 (e.g. JVASP-117419 LiCr2P2O8, JVASP-141543 Rb2LiFeF6) plus one commercial benchmark already in the set. For each, compute full ALIGNN-FF and DFT (auto-selected functional, same Lisv reference) delithiation curves on identical supercells; report Spearman rank correlation of vacancy orderings and MAE of step and average voltages. If vacancy-order Spearman falls below ~0.7 or average-voltage MAE exceeds ~0.4 V on more than half of the shortlist hosts, the primary screening signal is not reliable enough for the shortlists as stated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim rests on two pillars: (i) DFT-tier agreement with experiment on four commercial cathodes (≤0.3 V, ≤5% capacity) and (ii) the ALIGNN scalar head as a useful primary ranking signal (0.17 V MAE / R^{2}=0.94 vs ALIGNN-FF labels; 71 and 213 shortlists). Pillar (i) is supported for the four stoichiometric cases. Pillar (ii) is load-bearing for the shortlists and for the claim that single-pass ALIGNN is the primary screening signal (Abstract, Sec. 2.1–2.4, 2.9). The paper is explicit that the 0.17 V metric is distillation fidelity, not ALIGNN-vs-DFT (Sec. 2.1.2, 3.6). What remains least secure is whether ALIGNN-FF vacancy ranking and step voltages are close enough to DFT energetics that the distilled ranking remains useful for chemistries outside the five benchmarks. Sec. 2.6 already shows group-level rather than rank-level agreement and within-group order swaps (LFP/LMP, LMO/LCO); Sec. 2.8 notes that four of five commercial benchmarks are filtered by the Qgrav threshold and only LCO survives at rank 37. If ALIGNN-FF systematically mis-orders vacancies or misplaces plateaus for polyanion/fluoride hosts that dominate Table 2, the shortlists lose ranking value even though the DFT tier itself is sound. That is the softest condition for the central screening claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"BatteryMat is a three-tier hierarchical pipeline for average-voltage screening of Li-ion cathodes: (i) a single-pass ALIGNN scalar regressor trained on 7,610 ALIGNN-FF delithiation labels (MAE 0.17 V, R²=0.94 vs those labels), (ii) ALIGNN-FF step-by-step delithiation profiles, and (iii) automated supercell DFT with spacegroup-selected PBE+U or optB88-vdW+U and an in-house Li-metal reference that removes a ~1 V tabulated offset. On four stoichiometric commercial cathodes the DFT tier recovers experimental average voltage to within 0.3 V and theoretical volumetric capacity to within 5%; a non-stoichiometric NMC-like entry is treated as an edge case. The pipeline prioritises existing structures, returning 71 JARVIS-DFT and 213 Alexandria surrogate-level leads rather than confirmed new cathodes. Code and a web demo are released.","tokens_in":22711,"tokens_out":1640,"duration_ms":23913,"significance":"If the hierarchy is reliable, the work is a useful, reproducible methods contribution: it couples a carefully scoped ML front-end to a transparent DFT validation tier, with practical engineering choices (auto functional selection, recomputed μ_Li, inherited block-AFM MAGMOM) that other high-throughput cathode campaigns can adopt. Strengths include explicit scoping of the 0.17 V figure as distillation fidelity rather than DFT agreement, open code and AtomGPT web access, and a clean four-chemistry experimental benchmark. The paper does not claim generative discovery; its value is prioritisation at database scale with a documented validation path.","major_comments":[{"comment":"Sec. 2.1–2.2, 2.6, 2.9 and Table 2: The primary screening signal is a distillation of ALIGNN-FF labels. Sec. 2.6 already shows only group-level (not rank-level) agreement with DFT and within-group order swaps (LFP/LMP, LMO/LCO). The top-12 shortlist is dominated by polyanion phosphates and fluorides that are outside the five DFT benchmarks. Without at least a few DFT (or ALIGNN-FF vs DFT step-voltage) checks on those host families, the claim that single-pass ALIGNN is a useful primary ranking signal for the reported shortlists remains under-supported, even though the paper correctly labels them as surrogate-level leads.","section":"Sec. 2.1–2.2, 2.6, 2.9; Table 2"},{"comment":"Sec. 2.8: Four of the five commercial benchmarks are removed by the Q_grav,min = 20 mAh/g filter because JARVIS reduced-cell capacities are ~9–11 mAh/g; only LCO survives, at rank 37/71. A screen that systematically excludes known working cathodes under its default thresholds needs either a formula-unit renormalisation, a revised default threshold with a sensitivity analysis, or a clear demonstration that the composite score still recovers known chemistries when the capacity metric is put on a conventional scale. As written, this weakens confidence that the ranking protocol would surface useful candidates in practice.","section":"Sec. 2.8"},{"comment":"Sec. 2.4 vs Sec. 2.1: The Alexandria funnel derives surrogate voltages from ALIGNN formation-energy differences (Materials Project battery-explorer style) after ALIGNN-FF relaxation of charged hosts, whereas the JARVIS ALIGNN head is trained on sequential ALIGNN-FF vacancy-removal averages. These are different protocols. The manuscript should state explicitly whether the two shortlists are comparable, and whether the 0.17 V distillation metric applies to the Alexandria voltages at all (it does not, by construction).","section":"Sec. 2.4; Sec. 2.1"}],"minor_comments":[{"comment":"Abstract and Sec. 2.1: The phrasing ‘promotes single-pass average-voltage prediction with ALIGNN as the primary screening signal’ is easy to misread as ALIGNN-vs-experiment accuracy. Consider leading every abstract/results mention of 0.17 V with ‘vs ALIGNN-FF labels’ in the same sentence.","section":"Abstract; Sec. 2.1"},{"comment":"Table 1 / Sec. 2.5.4: NMC was run with PBE+U although the auto-selector would now route R-3m to optB88-vdW+U. Flagging this as future work is fine, but the table caption should state the functional mismatch more prominently so readers do not treat the +0.70 V residual as a pure framework failure.","section":"Table 1; Sec. 2.5.4"},{"comment":"Fig. 6 and Sec. 2.6: ALIGNN-FF volumetric capacities use unrelaxed JARVIS volumes while DFT uses relaxed CONTCAR volumes. A one-sentence note in the figure caption would prevent over-interpreting the capacity panel as a pure method comparison.","section":"Fig. 6; Sec. 2.6"},{"comment":"Sec. 2.1: ‘None of the five DFT-validated benchmark cathodes appear in the held-out test partition’ is good; also state that four fall in train and one in validation so readers do not assume full leave-out of commercial chemistries from the entire training process.","section":"Sec. 2.1"},{"comment":"Methods, lithium reference: Report the BCC Li lattice constant and k-mesh density used for μ_Li so the in-house values (−1.9031 / −0.9646 eV) are fully reproducible without reverse-engineering.","section":"Sec. 5"},{"comment":"Typos / notation: ‘V oltage’ with a space appears in several places (e.g. Introduction); ‘JV ASP’ inconsistently spaced; Eq. (1) uses μ_Li while the text sometimes writes μ Li. Standardise.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is carefully written and unusually honest about what the 0.17 V metric is and is not. I would not reject on novelty grounds: hierarchical ML+DFT cathode screening is incremental but the engineering package (auto functional, in-house μ_Li, open pipeline) is useful. The major-revision bar is the missing link between ALIGNN-FF ranking and DFT for the chemistries that actually dominate the shortlist, plus the Q_grav filter that drops known commercials. If the authors add even 2–3 DFT curves on top-ranked polyanion/fluoride hosts and fix or justify the capacity normalisation, this becomes a solid methods paper for a materials-informatics or computational-materials venue. Scope fit is good for cond-mat.mtrl-sci / computational materials journals; less so for a pure ML venue given the distillation framing."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is that this is a usable three-tier screening stack, not a new voltage theory. What is actually new is the coupling: single-pass ALIGNN as the front-end signal, ALIGNN-FF delithiation profiles, then automated supercell DFT with spacegroup-based functional routing (PBE+U vs optB88-vdW+U) and an in-house Li metal reference that removes the ~1 V tabulated offset. They ship code, a web demo, and two shortlists (71 from JARVIS-DFT, 213 from ~4.5 M Alexandria) that they correctly label as surrogate-level leads, not confirmed cathodes.\n\nWhat they do well is honesty about metrics. The 0.17 V MAE / R^{2}=0.94 is explicitly distillation fidelity to ALIGNN-FF labels, not agreement with DFT or experiment. On the four stoichiometric commercial cathodes the DFT tier recovers experimental average voltage within 0.3 V and theoretical volumetric capacity within 5%; the non-stoichiometric NMC-like entry is carried as an edge case with a transparent residual analysis (stoichiometry + oxygen redox). Methods detail (MAGMOM inheritance, supercell rule, cold-start restarts) is the kind of practical care that makes the pipeline reproducible.\n\nSoft spots, in proportion. The primary ranking signal is a force-field distillation; Sec. 2.6 already shows group-level rather than rank-level agreement and within-group order swaps, and four of five commercial benchmarks are filtered by the Qgrav threshold so only LCO survives the shortlist. That weakens how much weight you should put on the top-12 polyanion/fluoride list until more of those hosts go through DFT. The DFT validation set is five materials with no error bars; NMC was run with PBE+U before the auto-selector and is the obvious re-run. None of that breaks the central claim that the hierarchy is a practical way to prioritise existing database structures.\n\nThis is for computational battery groups and experimentalists who want a ranked shortlist with a clear path to DFT validation. Math and citations look standard and solid; free parameters (U values, windows, score weights) are stated. I would send it to peer review. Engage if you need a working screen or a clean Li-reference protocol; treat the shortlists as leads, not answers.","headline":"Solid, carefully scoped hierarchical cathode screen: DFT tier matches experiment on four commercial chemistries; the 0.17 V figure is distillation of ALIGNN-FF, not DFT, and the shortlists are openly surrogate leads.","tokens_in":23291,"tokens_out":590,"would_cite":true,"duration_ms":6260,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A three-tier machine-learning and DFT pipeline ranks lithium-ion cathode candidates by average voltage and recovers experimental voltages within 0.3 V on commercial chemistries.","keywords":["lithium-ion cathodes","high-throughput screening","graph neural networks","ALIGNN","DFT","average voltage","delithiation","JARVIS"],"falsifier":"Run the full DFT tier on the top-ranked non-commercial shortlist entries; if their computed average voltages deviate systematically by more than about 0.5 V from independent high-quality experiment or calculation while the four commercial benchmarks remain accurate, the surrogate primary filter has failed.","tokens_in":23031,"feed_emoji":"🔋","tokens_out":1083,"duration_ms":24460,"temperature":0.7,"pith_summary":"Average intercalation voltage is the primary figure of merit for cathode screening, yet full DFT delithiation curves cannot scale to modern materials databases, while pure machine-learning surrogates skip thermodynamic consistency. BatteryMat places a single-pass graph-network voltage predictor at the front of a three-tier cascade, then validates survivors with force-field delithiation profiles and automated supercell DFT that chooses the exchange-correlation functional by spacegroup and recomputes the lithium-metal reference in the same plane-wave basis. On four stoichiometric commercial cathodes the DFT tier matches experimental average voltage to within 0.3 V and theoretical volumetric capacity to within 5 percent. Applied to existing databases the pipeline returns 71 lithium candidates from JARVIS-DFT and 213 from roughly 4.49 million Alexandria structures—all prioritised surrogate-level leads rather than confirmed new materials. The practical claim is that average voltage can drive high-throughput ranking without sacrificing experimental agreement at the validation end.","feed_headline":"Three-tier screen recovers cathode voltages within 0.3 V","feed_subtitle":"Graph-network ranking plus force fields and DFT prioritises leads from millions of existing structures.","key_machinery":"The three-tier hierarchy: an ALIGNN scalar average-voltage predictor that collapses an N-step force-field protocol into one forward pass (0.17 V MAE against its ALIGNN-FF labels), an ALIGNN-FF step-by-step delithiation tier that ranks lithium vacancies, and automated supercell DFT with automatic functional selection and a recomputed lithium chemical potential.","core_discovery":"BatteryMat establishes that a single-pass ALIGNN regressor, trained on 7,610 force-field delithiation voltages, can serve as the primary screening signal for lithium-ion cathodes; when survivors are advanced through ALIGNN-FF profiles and automated PBE+U or optB88-vdW+U supercell DFT (with spacegroup-selected functionals and an in-house lithium reference that removes a roughly 1 V tabulated offset), the pipeline recovers experimental average voltages to within 0.3 V and crystallographic volumetric capacities to within 5 percent on four commercial chemistries, while ranking existing database structures into shortlists of prioritised candidates rather than inventing new ones.","pith_inferences":["Ensembling several universal force fields at the second tier would turn vacancy-ranking disagreement into an explicit confidence band before DFT is spent.","Retraining the scalar head on true DFT voltages instead of force-field labels would convert the 0.17 V figure into a direct DFT-distillation error and change how users interpret the primary screen.","Once a few sodium and magnesium DFT anchors exist, the same hierarchy can be re-ranked for multivalent hosts, because the surrogate already recovers the expected voltage–capacity trade-off for those ions.","Finite-supercell staircases on two-phase materials such as LiFePO4 imply that larger cells or explicit phase-boundary models will be required before the method can claim plateau-shape fidelity rather than only average-voltage fidelity."],"forward_implications":["Database-scale cathode ranking becomes feasible in minutes on a single GPU before any DFT budget is spent.","Layered frameworks are automatically routed to optB88-vdW+U and 3D-bonded frameworks to PBE+U without per-material manual choice.","Recomputing the lithium-metal reference in the cathode plane-wave basis removes a systematic ~1 V offset and places voltages on a common experimental scale.","The 71 JARVIS and 213 Alexandria shortlists become the concrete input for the next round of DFT validation campaigns.","For the four chemistries with DFT residuals below 0.3 V, cell-level energy density estimated as V_avg times theoretical capacity stays within roughly 8 percent of experiment."],"fun_headline_variants":["BatteryMat: ALIGNN then DFT recovers Li cathode voltages within 0.3 V","Three-tier screen ranks millions of structures for cathode voltages","ALIGNN-FF-DFT pipeline matches experimental voltages to 0.3 V","Hierarchical GNN plus DFT prioritises existing Li-ion cathode candidates","BatteryMat ranks JARVIS and Alexandria pools for average-voltage leads"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The cascade assumes that the force-field delithiation protocol ranks lithium vacancies and plateaus faithfully enough that a single neural-network pass distilled from those labels still orders real cathode candidates usefully.","fun_headline_variants_meta":{"raw":{"variants":["BatteryMat: ALIGNN then DFT recovers Li cathode voltages within 0.3 V","Three-tier screen ranks millions of structures for cathode voltages","ALIGNN-FF-DFT pipeline matches experimental voltages to 0.3 V","Hierarchical GNN plus DFT prioritises existing Li-ion cathode candidates","BatteryMat ranks JARVIS and Alexandria pools for average-voltage leads"]},"model":"grok-4.5","effort":"low","cost_usd":0.005202,"raw_usage":{"total_tokens":1585,"prompt_tokens":981,"num_sources_used":0,"completion_tokens":79,"cost_in_usd_ticks":52020000,"prompt_tokens_details":{"text_tokens":981,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":525,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":981,"tokens_out":79,"duration_ms":5984,"temperature":1.0,"reasoning_tokens":525,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T00:36:32.695367+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the full DFT tier on the top-ranked non-commercial shortlist entries; if their computed average voltages deviate systematically by more than about 0.5 V from independent high-quality experiment or calculation while the four commercial benchmarks remain accurate, the surrogate primary filter has failed.","supporting_citations":[],"review_version":1}