{"id":"05b009a8-810b-4486-9abb-3fdac6e0c296","arxiv_id":"2607.27430","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Independent CNNs plus an Alice–Bob uncertainty pair extract Aex, Ku, D, and Npin from FORC images and recreate the input magnetometry in closed loop.","lead":"CNNs can recover the full micromagnetic Hamiltonian (exchange, anisotropy, DMI, pinning) from FORC magnetometry images. Closed-loop resimulations and an Alice–Bob uncertainty network make the inverse problem usable on both simulated and experimental thin-film data.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Experimental closed-loop is not a pure CNN inversion: Ms is swept and pinning is ensemble-averaged before visual selection.","rationale":"The reader correctly flags the simplified zero-T Hamiltonian plus ad-hoc pinning as the soft spot for experimental transfer and already rates the work CONDITIONAL proof-of-concept. The more precise load-bearing gap is methodological: the experimental “closed-loop” is not the same test as the simulated one (F_sim∘F_CNN), because Ms sweeping and pinning-ensemble averaging plus visual selection inflate apparent success. That does not overturn the simulated results or the Alice–Bob idea, so the verdict stays CONDITIONAL rather than moving to REJECT; it simply tightens why experimental support is only qualitative. No code/data release still limits re-running. I agree with the reader’s neighborhood of concern but locate the weakest link in the experimental validation protocol itself, not only in model misspecification.","tokens_in":12173,"tokens_out":570,"duration_ms":36080,"concrete_test":"Freeze Ms at the volume-normalized multilayer value (~560 kA/m) or bulk-Co (1430 kA/m) with no sweep; take the single CNN output (Npin,Aex,Ku,D); simulate without averaging over pinning ensembles (or report the full distribution); score quantitative FORC error (e.g. MAE on the 64×64 image, plus ΔHc and loop-area error) against the VSM data. If error rises sharply versus the hand-picked Ms=1307 case, the experimental closed-loop does not validate pure CNN inversion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim requires that predicted Hamiltonians recreate experimental FORCs via closed-loop verification. In Sec. III.C the protocol is: (i) treat Ms as free and sweep [470,1400] kA/m; (ii) for each Ms run the CNNs; (iii) average mumax3 FORCs over 20 random pinning realizations at the predicted Npin; (iv) report the Ms (1307 kA/m) with best visual match. That procedure adds free parameters and post-hoc selection outside the trained inverse map F_CNN. Fig. 5(d) therefore shows that some nearby parameter set can roughly recover coercivity/squareness, not that the CNNs alone decode the experimental Hamiltonian from the FORC image. The paper already notes poorer agreement than simulated closed-loops and missing higher-order/microstructural terms—so the experimental half of the claim is weaker than the abstract states. Simulated closed-loops (Fig. 4) remain supportive inside the training model class, especially for feature-rich FORCs.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The manuscript trains a set of independent CNNs to invert First-Order Reversal Curve (FORC) images for the phenomenological micromagnetic parameters Aex, Ku, D, and Npin (with Ms injected as a known macroscopic input). Training and primary validation use zero-temperature mumax3 simulations that include ad-hoc high-Ku pinning cells and polycrystalline anisotropy patches. Performance is reported via Pearson correlations and mean absolute errors on a held-out validation set, a closed-loop re-simulation of ten validation FORCs, an Alice–Bob dual network that predicts Alice’s error quantile from FORC features alone, and a qualitative comparison to experimental Co/Pd multilayer FORCs after an Ms sweep and ensemble averaging over pinning realizations.","tokens_in":12447,"tokens_out":1393,"duration_ms":35740,"significance":"If the inverse map is reliable inside and near the training distribution, the work offers a practical, data-driven alternative to iterative trial-and-error fitting of micromagnetic Hamiltonians from readily measured FORCs. Strengths that should be credited include: (i) decoupled per-parameter CNNs, which make the claim that each term leaves an independent fingerprint more stringent; (ii) explicit closed-loop re-simulation on feature-rich synthetic FORCs (Fig. 4); and (iii) the Alice–Bob construction, which estimates uncertainty without ground-truth parameters. These elements are concrete and falsifiable within the stated model class. The experimental half of the claim and the generality beyond the simplified disorder model are weaker and currently limit the impact relative to the abstract’s wording.","major_comments":[{"comment":"Abstract and Sec. III.C claim closed-loop recreation of experimental FORCs by the CNN-predicted Hamiltonian. The actual protocol (i) treats Ms as a free parameter swept over [470, 1400] kA/m, (ii) averages mumax3 FORCs over 20 random pinning profiles at the predicted Npin, and (iii) selects the Ms (1307 kA/m) by best visual match. That procedure introduces free parameters and post-hoc selection outside the trained map F_CNN. Fig. 5(d) therefore shows that some nearby parameter set can roughly recover coercivity/squareness, not that the CNNs alone decode the experimental Hamiltonian from the FORC image. The abstract and experimental-validation claims should be revised to state this protocol explicitly and to qualify the experimental result as qualitative consistency under Ms/pinning post-processing, not pure closed-loop inversion.","section":"Abstract; Sec. III.C; Fig. 5(d)"},{"comment":"Primary training and closed-loop checks (Fig. 4) use the same mumax3 forward model and the same simplified disorder (Npin cells with Ku = 50 MJ/m³ plus polycrystalline patches). The paper itself notes that experimental agreement is poorer than the best simulated cases and that higher-order terms and microstructural detail are missing (Sec. III.C, Conclusion). This is a mild self-consistency loop for the synthetic half of the claim. To keep the central claim load-bearing, the manuscript should either (a) quantify how far experimental FORCs lie from the training distribution (e.g., via Bob’s uncertainty or a domain-shift metric) or (b) add at least one held-out synthetic test with a disorder model not used in training (different pin strength, grain-size distribution, or finite-T nucleation) so that generalization beyond the training Hamiltonian class is demonstrated rather than asserted.","section":"Sec. II.B; Sec. III.A; Sec. III.C; Conclusion"},{"comment":"Fig. 4 and the accompanying text show that featureless FORCs (IDs 768, 1503, 2529)—those that collapse onto the major loop—are poorly reconstructed. The text correctly attributes information content to minor-loop structure, but the abstract and conclusion still present the method as extracting the “full” Hamiltonian from FORCs in general. The applicability domain should be stated up front (feature-rich minor loops required) and, if possible, Bob’s predicted uncertainty should be shown to flag precisely these featureless cases on the validation set, turning the acknowledged failure mode into a usable rejection criterion.","section":"Sec. III.A; Fig. 4; Abstract; Conclusion"}],"minor_comments":[{"comment":"Total number of simulated samples that passed screening and entered the 80/20 split is never stated; only a ~20% success rate is given. Report N_train and N_val explicitly.","section":"Sec. II.B–C"},{"comment":"Table I lists D max = 0.005 J/m²; units and physical range should be cross-checked against typical interfacial DMI values and the cell size used.","section":"Table I"},{"comment":"Fig. 1(b) and the 64×64 rearrangement are clear, but the precise field-sampling and quantization procedure (byte scaling of mz) should be stated so that experimental FORCs can be mapped identically.","section":"Sec. II.A; Fig. 1"},{"comment":"Typos / wording: “microsctuctures” (Introduction); “decoupled parameter prediction” is used without defining the training independence earlier; “close-loop” vs “closed-loop” inconsistency in figure captions.","section":"Introduction; Sec. III.A; Fig. 4 caption"},{"comment":"Alice–Bob: state Bob’s architecture (same CNN backbone or different), loss for quantile regression, and whether Bob is trained on the validation-set error distribution or a nested split, to avoid leakage into the uncertainty calibration.","section":"Sec. III.B; Fig. 5(a)"}],"recommendation":"major_revision","confidential_remarks":"The synthetic closed-loop and Alice–Bob ideas are publishable and appropriate for a materials/ML journal. The main risk is over-claiming on the experimental side; if the authors tone the abstract and Sec. III.C to match what Fig. 5(d) actually shows and add a clear applicability statement for featureless FORCs, the paper becomes a solid methods contribution. I would not reject on novelty or on the use of a simplified disorder model alone."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real advance here is straightforward: they take the earlier DMI-from-FORC CNN and scale it to four independent networks that recover Aex, Ku, D, and Npin from a 64×64 FORC image (plus injected Ms). Closed-loop resimulation on the validation set mostly works when the FORCs have structure; featureless abrupt-switch cases fail as expected. Alice–Bob is the cleanest piece—Bob maps the same image to Alice’s error quantile without ever seeing ground truth, and the binned plots show it actually separates the high-error tail.\n\nInside the mumax3 training distribution the numbers are fine (Pearson 0.84–0.97, MAE 0.04–0.11). The independent-network design is a plus: no cheating via cross-parameter correlations. They are also candid that the information lives in the minor-loop deviations, not the major loop.\n\nSoft spots are real but proportionate. Training and primary validation sit inside the same zero-T, ad-hoc-pinning forward model, so the simulated closed loops are partly self-consistency. The experimental Co/Pd claim is softer still: they sweep Ms over a wide range, average 20 pinning realizations, and pick the best visual match. That is not a pure CNN inversion; Fig. 5(d) shows a nearby parameter set can roughly recover coercivity and squareness, not that the networks alone decoded the film. The paper itself notes the poorer agreement and missing higher-order/microstructural terms, so the abstract over-reaches a bit on “experimental FORCs.” No code or data release limits immediate checks.\n\nThis is for people who already run FORCs or micromagnetic inverse problems and want a practical starting Hamiltonian plus an uncertainty flag. It will not replace multi-technique campaigns yet, but it is a usable proof-of-concept method paper. I would send it to referees; the core idea and the Alice–Bob construction deserve the scrutiny. Worth a look if you touch thin-film magnetometry or ML-for-spintronics; skip if you need quantitative experimental Hamiltonians tomorrow.","headline":"Solid extension of their DMI-FORC CNN to a full Hamiltonian inverse map, with honest closed-loop checks and a useful no-GT uncertainty net; experimental half is weaker than the abstract.","tokens_in":13074,"tokens_out":540,"would_cite":true,"duration_ms":17174,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Convolutional networks can extract the full micromagnetic Hamiltonian directly from First-Order Reversal Curve fingerprints.","keywords":["micromagnetic Hamiltonian","First-Order Reversal Curves","convolutional neural networks","Dzyaloshinskii-Moriya interaction","magnetometry","uncertainty quantification","closed-loop verification"],"falsifier":"Measure FORCs on a film whose exchange, anisotropy and DMI are already known by independent techniques (Brillouin light scattering, ferromagnetic resonance, domain-wall creep); if the networks return values far outside those bounds and the re-simulated FORC fails to match the measured one, the claim is falsified.","tokens_in":13043,"feed_emoji":"🧲","tokens_out":851,"duration_ms":31655,"temperature":0.7,"pith_summary":"Macroscopic magnetometry averages over every spin, so recovering the underlying micromagnetic Hamiltonian has long looked underdetermined. This paper shows that the detailed shapes of First-Order Reversal Curves still carry enough independent fingerprints of exchange, anisotropy, Dzyaloshinskii–Moriya interaction and pinning density that a set of convolutional networks can invert them. The networks are trained only on simulated FORCs; when their predicted parameters are fed back into the same micromagnetic simulator, the original curves are recovered for held-out simulations and, to a useful degree, for an experimental Co/Pd multilayer. A second parallel network estimates how trustworthy each prediction is from the FORC shape alone, without ever seeing ground-truth labels. If the approach holds, routine vibrating-sample-magnetometer data can replace slow trial-and-error fitting and specialized spectroscopies when modeling complex magnetic materials.","feed_headline":"CNNs extract full spin Hamiltonians from FORC fingerprints","feed_subtitle":"Predicted parameters re-simulate the original magnetometry; a twin network scores confidence without ground truth","key_machinery":"Independently trained CNNs that invert FORC images into each Hamiltonian parameter, verified by closed-loop mumax3 re-simulation and accompanied by an Alice–Bob dual network that predicts the quantile of Alice’s error without ground-truth labels.","core_discovery":"A collection of independently trained convolutional neural networks maps 64-by-64 FORC images (plus the measured saturation magnetization) onto the full phenomenological micromagnetic Hamiltonian—exchange stiffness, uniaxial anisotropy, Dzyaloshinskii–Moriya interaction and pinning-site density—while a parallel “Alice–Bob” network quantifies prediction uncertainty from the same FORC input alone. Closed-loop re-simulation with the predicted parameters recreates the input magnetometry for both synthetic validation samples and an experimental Co/Pd film.","pith_inferences":["The visibly poorer experimental match versus the best simulated closed-loops implies that residual microstructure still leaves identifiable fingerprints that larger, more realistic training sets could capture.","Because each parameter is predicted by a strictly independent network, the method can test which energy terms remain separately observable after ensemble averaging.","Alice–Bob uncertainty maps offer a practical pre-filter for high-throughput materials libraries: only high-confidence FORCs need expensive follow-up characterization."],"forward_implications":["Routine VSM FORC measurements can supply quantitative Hamiltonian parameters without specialized spectroscopies.","Prediction confidence can be scored from the FORC shape alone, flagging featureless or ambiguous samples before any further work.","Closed-loop re-simulation becomes a standard check that extracted parameters actually reproduce the observed magnetometry.","The same workflow extends immediately to richer Hamiltonians once training sets include higher-order interactions and realistic microstructure."],"fun_headline_variants":["CNNs decode full micromagnetic Hamiltonian from FORC fingerprints","Neural nets map 64x64 FORCs to exchange, anisotropy, DMI, pinning","Alice-Bob CNNs extract spin Hamiltonian and score uncertainty solo","Predicted Hamiltonian params re-simulate original FORC magnetometry","CNNs recover phenomenological spin Hamiltonian directly from FORCs"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The simplified zero-temperature micromagnetic model with artificial high-anisotropy pinning cells and polycrystalline patches is rich enough that a network trained only on it can correctly invert real experimental FORCs.","fun_headline_variants_meta":{"raw":{"variants":["CNNs decode full micromagnetic Hamiltonian from FORC fingerprints","Neural nets map 64x64 FORCs to exchange, anisotropy, DMI, pinning","Alice-Bob CNNs extract spin Hamiltonian and score uncertainty solo","Predicted Hamiltonian params re-simulate original FORC magnetometry","CNNs recover phenomenological spin Hamiltonian directly from FORCs"]},"model":"grok-4.5","effort":"low","cost_usd":0.004503,"raw_usage":{"total_tokens":1291,"prompt_tokens":702,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":45028000,"prompt_tokens_details":{"text_tokens":702,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":518,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":702,"tokens_out":71,"duration_ms":8539,"temperature":1.0,"reasoning_tokens":518,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T01:07:22.765391+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Measure FORCs on a film whose exchange, anisotropy and DMI are already known by independent techniques (Brillouin light scattering, ferromagnetic resonance, domain-wall creep); if the networks return values far outside those bounds and the re-simulated FORC fails to match the measured one, the claim is falsified.","supporting_citations":[],"review_version":1}