{"id":"57f6ebde-5aa8-431f-9f27-3babbca03305","arxiv_id":"1908.09883","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Transfer learning with (L,2)-tiling, which repeats a small-system weight pattern into a larger restricted Boltzmann machine, reaches the ground state faster and more accurately than random initialization in several quantum phases.","lead":"The authors test whether reusing a neural-network quantum state trained for a small spin system can speed up and improve calculations for larger systems. They introduce several 'tiling' schemes to copy weights from small to large restricted Boltzmann machines and show that, for some quantum phases, this beats starting from random weights.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'more effective' half of the headline claim lacks a direct cold-start baseline in the energy-error plots; Figs. 7 and 11 compare transfer protocols only against MPS/QMC references, not against random initialization.","rationale":"The paper's stated central claim is comparative: transfer learning should beat cold-start in both efficiency and effectiveness. Efficiency is directly benchmarked against cold-start in Fig. 4, but effectiveness is only benchmarked among transfer protocols against external references. The text asserts that cold-start is trapped in a local minimum for the Heisenberg model, yet no cold-start energy error is reported. This is a load-bearing gap because a protocol can be fast but not more effective; the compound claim requires both comparisons. The reader's verdict is already CONDITIONAL and asks for fuller accounting of excluded runs, which is related but not identical to the missing cold-start effectiveness baseline I identify. I therefore do not propose changing the verdict: the conditional status already captures the need for additional evidence. My concern does not accuse the authors of manipulation; it simply notes that one half of the headline comparison is currently asserted rather than displayed.","tokens_in":17541,"tokens_out":14567,"duration_ms":161811,"concrete_test":"Re-run the 20 realizations for cold-start and each tiling at L=128 (and 8x8) for the same models and phases, and plot the energy relative error of cold-start alongside the transfer protocols at identical wall-clock time budgets (e.g., the total accumulated time at which (L,2)-tiling reaches its stopping criterion). If cold-start's mean and percentile energy error is not larger than that of the best tiling in the phases where the paper claims effectiveness, the 'far more effective' part of the abstract fails; if cold-start errors are substantially larger, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that some tilings are 'far more effective and efficient' than cold-start. Fig. 4 directly supports the efficiency half (time to stopping criterion vs cold-start), but the effectiveness plots do not include cold-start. Fig. 7(a,b) and Fig. 11(b) report energy relative errors for the (L,2), (1,2), and (2,2) tilings relative to MPS/QMC references only; cold-start accuracy is never shown. The only quantitative cold-start effectiveness statement in Sec. VI B is for the 4->128 Ising scenario ('30% better than the cold-start'), a single case. For Heisenberg, the paper asserts that cold-start is 'trapped in a local minimum prematurely' but provides no energy error for that baseline. Since 'effective and efficient' is a compound claim, a missing effectiveness baseline means the headline is under-supported even if all transfer protocols are accurate in absolute terms. The issue is compounded by the explicit exclusion of non-converged (1,2)/(2,2) Heisenberg realizations from Fig. 7(b), which makes the transfer-vs-transfer comparison itself dependent on an unreported filtering rule.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes transfer learning protocols for restricted-Boltzmann-machine neural-network quantum states, in which weights trained for a small system are replicated (with prescribed (k,p)-tilings) to initialize a larger system. The protocols are benchmarked on the transverse-field Ising and Heisenberg XXZ models in one dimension and on the latter in two dimensions, with system sizes up to 128 and 8x8 spins and 20 realizations per setting. Efficiency is measured as time to a variance-based stopping criterion; effectiveness is measured by energy and spin-correlation errors relative to matrix-product-state or quantum Monte Carlo references. The paper reports that one protocol, (L,2)-tiling, is more efficient and effective than cold-start random initialization for the Ising model, that transfer protocols can avoid premature local minima for the Heisenberg model, and that a TensorFlow/GPU implementation is faster than the NetKet CPU implementation in comparable settings.","tokens_in":17635,"tokens_out":7432,"duration_ms":77508,"significance":"If the central claim is fully supported, this is a useful empirical contribution to scaling neural-network quantum states. The study is extensive in scope: two model families, multiple phases, 1D and 2D, several system sizes up to 128 and 8x8, 20 realizations, and external MPS/QMC references. The explicit transfer-distance diagnostic is a valuable idea, as is the physics-motivated distinction among tilings that preserve or destroy the relevant magnetic correlations. The TensorFlow port with GPU speedup is a practical contribution of independent interest. However, the headline claim of being 'far more effective and efficient' than cold-start is not currently backed by a direct cold-start effectiveness baseline in the main energy-error figures, and one figure excludes non-converged realizations without reporting exclusion counts. These gaps are fixable within the scope of the manuscript and do not, on the evidence presented, invalidate the underlying approach.","major_comments":[{"comment":"The abstract and conclusions claim that some transfer protocols are 'far more effective and efficient' than cold-start random initialization, and Sec. IV states that the evaluation will compare transfer-trained networks with networks trained from random initialization. The efficiency half of the claim is directly supported by Fig. 4, but the effectiveness half is not: Figs. 7 and 11(b) report energy errors of the transfer protocols only against MPS/QMC references, with no cold-start energy error plotted. The only quantitative cold-start effectiveness comparison is the 4-to-128 Ising scenario in Sec. VI.B and Fig. 10, where the transfer energy error is described as 30% better than cold-start. Please add cold-start energy-error baselines to the same effectiveness evaluations (at the same iteration counts used in Fig. 7), or explicitly restrict the claim to efficiency plus the specific scenarios where cold-start effectiveness has been measured.","section":"§VI.B, Figs. 7 and 11"},{"comment":"The text states that for the Heisenberg model with Δ=-0.5 and Δ=-2.0, 'some realizations' of the (1,2)- and (2,2)-tiling protocols fail to converge because of large gradients, and that these extreme cases are not included in the error bars of Fig. 7(b). The number of excluded realizations is not reported for any protocol, phase, or size. Because these excluded runs belong to the very protocols against which (L,2)-tiling is compared, the reported error bars may systematically understate the failure rate and variance of those protocols. Please report the exclusion count for every affected entry and provide a sensitivity analysis, such as a conservative worst-case treatment of the non-converged realizations.","section":"§VI.B, Fig. 7(b)"},{"comment":"The three transfer protocols are all initialized from the same base network, namely the solution of the 'most effective protocol' at Lv=64. This common-base choice is intended to make the comparison fair, but it may instead disadvantage the (1,2)- and (2,2)-tilings, whose weight-copying rules are designed to be applied to base weights produced by those same protocols. A base trained under (L,2)-tiling will have a weight structure that the (1,2)- and (2,2)-tilings may not preserve. Please either use each protocol's own base network at Lv=64 for the Lv=128 evaluation, or justify why a common base is the appropriate comparison and report the protocol that supplied the base for each panel.","section":"§VI.B, Fig. 7 and caption"},{"comment":"The effectiveness results in Fig. 7 are evaluated at a fixed number of iterations determined by the stopping criterion of the fastest transfer protocol, not at each protocol's own stopping point. This makes the comparison a time-limited or iteration-limited accuracy comparison rather than a comparison of final converged states, which is how 'effectiveness' is defined in Sec. V. The paper should state this explicitly in the main text and in the figure captions, and should also include the cold-start protocol in the same fixed-iteration evaluation. Otherwise the reader cannot separate accuracy at equal computational budget from accuracy at convergence, which matters for the paper's claim that transfer learning improves scalability.","section":"§V.B and §VI.B, definition of effectiveness evaluation"}],"minor_comments":[{"comment":"The correlators CF_d and CA_d are defined with a factor 1/(d-1), but the text and figures evaluate them for d starting at 1, for which the denominator vanishes. Please either start the range at d=2 or define the d=1 case separately.","section":"§V.A, Eqs. (12) and (13)"},{"comment":"The index order in Eq. (14) is inconsistent with the convention W_ji used earlier in the paper: Eq. (14) writes W_{i,j}, while the surrounding text and Fig. 5 use W_{j,i}. Please make the indices consistent.","section":"§V.A, Eq. (14)"},{"comment":"The text says that a hidden configuration is sampled from p(h|x) using Eq. (7) and that a visible configuration is sampled from p(h|x) using Eq. (8); the two formulas appear to be labeled in the opposite order. Please correct the labeling or the description.","section":"§II, Eqs. (7) and (8)"},{"comment":"The parameters of the optimization are said to be determined from the literature or from a grid search, but no grid, range, or selection criterion is given. For reproducibility, please list the grid-search choices and the selected values.","section":"§V.B"},{"comment":"The statement that simulations with biases set to zero and with variable biases yield very close results is not accompanied by data. Please show a comparison or provide a reference to a figure/table supporting this claim.","section":"§IV"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a timely and practical question, and the empirical scope is substantial. The main barrier is not the idea but the support for the headline effectiveness claim: the principal energy-error figures lack a cold-start baseline, the exclusion of non-converged runs is undocumented, and the common-base choice in Fig. 7 deserves justification. These are fixable with additional analysis and reporting. The paper could also be strengthened by a code/data availability statement, since the study is empirical."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The (k,p)-tiling rules are new and physically motivated: (L,2)-tiling preserves antiferromagnetic correlations, (1,2)-tiling biases toward ferromagnetic order, and the paper tests these ideas seriously across phases, dimensions, and models. The efficiency half of the central claim is solid: Fig. 4 shows (L,2)-tiling beats cold-start time to the stopping criterion for Ising in all phases and for Heisenberg in the AFM and XY phases. The effectiveness half is under-supported. Figures 7 and 11 plot transfer-protocol energy errors against MPS/QMC references, but neither includes a cold-start baseline. The only quantitative cold-start effectiveness number in the paper is the 4->128 Ising case (30% better error), and for Heisenberg the cold-start trap is asserted, not shown. So the abstract's 'far more effective and efficient' exceeds what the reported data demonstrate.\n\nGood things: the evaluation is broad—twenty realizations, 1D and 2D up to 128 and 8x8, external references, phase-resolved. The transfer-distance analysis is simple and informative. The work is clean on the circularity axis: accuracy is judged against MPS/QMC, and no claim reduces to a fitted parameter. The GPU port for NetKet-style NQS is a practical contribution, though code and data are not released.\n\nSoft spots, in proportion. (1) The missing cold-start effectiveness baseline is the main issue; it is fixable and should be fixed. (2) Non-converged Heisenberg realizations for (1,2)- and (2,2)-tilings are excluded from Fig. 7(b) without reporting how many. That can bias protocol comparisons; the authors need to report counts and show sensitivity. (3) The success of a tiling depends on knowing the phase in advance; the authors acknowledge this, but it should be visible in the abstract. (4) No code/data release is a minor reproducibility gap for the GPU speedup claim.\n\nThe paper is a serious empirical method study with a genuine new idea and mostly honest execution. For anyone working on NQS scalability, it is worth reading and citing. It deserves peer review. My recommendation: send to a reputable journal, but require major revision so the authors add cold-start effectiveness baselines and a full accounting of excluded runs. The efficiency result alone may not justify the current headline.","headline":"Transfer learning with (k,p)-tilings is a genuinely useful new trick for scaling NQS, but the headline 'more effective than cold-start' needs a cold-start effectiveness baseline that the current figures don't show.","tokens_in":18323,"tokens_out":5369,"would_cite":true,"duration_ms":49433,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Trained neural-network quantum states can be reused to jump-start ground-state searches on larger spin systems, reaching the target faster and more accurately than a random start when the transferred pattern matches the magnetic phase.","keywords":["neural-network quantum states","restricted Boltzmann machine","transfer learning","variational Monte Carlo","transverse-field Ising model","Heisenberg XXZ model","quantum phases","GPU implementation"],"falsifier":"On the one-dimensional transverse-field Ising chain with $J_I = -2$ (antiferromagnetic phase), scale from 4 to 128 spins using the $(1,2)$-tiling, $(L,2)$-tiling, and a cold start with the paper's stopping criterion; the central claim fails if the $(L,2)$-tiling hot start does not reach the stopping criterion in less time and with lower mean relative energy error than the cold start over the 20 realizations.","tokens_in":17169,"feed_emoji":"⚛️","tokens_out":9198,"duration_ms":91726,"temperature":0.7,"pith_summary":"The paper asks whether a neural-network quantum state trained on a small spin system can be reused to initialize the search for the ground state of a larger version of the same model. It proposes several weight-transfer protocols, called tilings, and tests them on one- and two-dimensional Ising and Heisenberg models. The central claim is that some protocols, especially the (L,2)-tiling, reach the stopping criterion faster and with lower energy error than starting from random weights, provided the tiling preserves the correlation pattern of the target quantum phase. If true, transfer learning gives a practical route to scaling variational many-body simulations to larger system sizes and reduces the risk of getting trapped in local minima. The catch is that the phase must be known in advance: a mismatched tiling can be slower and less accurate than a cold start.","feed_headline":"Phase-matched warm starts speed up neural-network quantum states","feed_subtitle":"Small-lattice weights reach larger ground states faster—if the tiling preserves the magnetic order.","key_machinery":"The central object is the restricted Boltzmann machine (RBM), a two-layer probabilistic neural network whose visible nodes encode spin configurations and whose squared amplitude defines the variational wave function. The transfer mechanism is the (k,p)-tiling rule: copy groups of k rows of the trained weight matrix and repeat each group p times across the target network's hidden nodes, filling the remaining weights with small random values. This mapping is what carries the learned correlations from the base system to the target; it succeeds only when the repeated tile matches the target phase's spin-spin correlation pattern. The implementation also ports the RBM energy minimization to GPU hardware, which is what makes the timing comparisons meaningful.","core_discovery":"The paper's central claim is that transfer learning can improve the scalability of neural-network quantum states: a restricted Boltzmann machine trained on a small lattice can initialize the same ansatz on a larger lattice, and with the right tiling this hot start reaches the stopping criterion faster and with lower error than a cold start. The (L,2)-tiling protocol, which repeats the entire learned weight matrix in blocks, is the best all-around protocol for the one-dimensional transverse-field Ising model in all phases tested, and is competitive for the Heisenberg XXZ model in the antiferromagnetic and XY phases. The (1,2)- and (2,2)-tilings are better for the ferromagnetic Heisenberg phase, where magnetization is conserved, because they preserve the expected domain-wall structure. In the mismatched case, such as the (1,2)-tiling in the antiferromagnetic phase, the hot start can be slower and less accurate than a cold start, so the phase must be known in advance.","pith_inferences":["The phase-dependence of the best tiling suggests a practical selector that estimates the target's dominant correlations from a short cold start or a classical pre-solve and then chooses the matching tiling; the paper does not test such a selector.","The same tiling idea should transfer between models that share a phase, for example from the antiferromagnetic Ising chain to the antiferromagnetic Heisenberg chain, because the correlation pattern the tiling preserves is the same; the paper lists cross-Hamiltonian transfer as future work rather than a result.","For two-dimensional systems with striped or other non-uniform order, the isotropic (L,2)-tiling is only a special case; per-axis repetition factors could yield better performance, but this is not explored here.","The reported non-convergent Heisenberg realizations after a bad transfer initialization imply that transfer learning changes the structure of the optimization landscape, so a transfer-distance diagnostic may need to be supplemented by a convergence-prediction criterion."],"forward_implications":["For one-dimensional transverse-field Ising chains, the (L,2)-tiling hot start reaches the stopping criterion faster and with lower mean relative energy error than a cold start in ferromagnetic, antiferromagnetic, and paramagnetic phases.","For the Heisenberg XXZ chain at fixed zero magnetization, a cold start reaches the stopping criterion fastest but can be trapped in a local minimum, whereas the best phase-matched tiling gives lower energy error, except in the ferromagnetic phase where the (L,2)-tiling performs worse.","Scaling a one-dimensional Ising chain from 4 to 128 spins in a single transfer already beats a cold start, and larger jumps early in the transfer followed by smaller jumps are generally better than uniform doubling.","The spin-spin correlators produced by the best phase-matched tiling track the matrix-product-state reference values more closely than mismatched tilings, so the accuracy gain is not limited to the energy.","In the two-dimensional antiferromagnetic Heisenberg lattice, the (L,2)-tiling is the most efficient and effective of the three protocols tested."],"supporting_citations":[{"why":"It introduces the restricted-Boltzmann-machine ansatz whose trained weights are transferred.","marker":"[11]"},{"why":"It supplies the transfer-learning framing of reusing a trained model for a related task.","marker":"[20]"},{"why":"It supplies evidence that features transfer between same-architecture neural networks, motivating the hot-start initialization.","marker":"[22]"},{"why":"It identifies the open-source CPU implementation that the GPU port is compared against.","marker":"[19]"},{"why":"It is the specific CPU implementation used as the timing baseline in the speed comparison.","marker":"[26]"},{"why":"It provides the GPU machine-learning platform on which the implementation runs.","marker":"[21]"},{"why":"It supplies the matrix-product-state reference energies and correlators for the one-dimensional models.","marker":"[6]"},{"why":"It supplies the quantum Monte Carlo reference energies and correlators for the two-dimensional model.","marker":"[2]"}],"fun_headline_variants":["Warm starts from small lattices beat random init","Blockwise tiling fastest for Ising ground states","Transfer learning scales neural quantum states","Reused weights cut time to larger ground states","Phase-aware tiling accelerates NQS training"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the target system's magnetic phase is known well enough to choose a tiling that preserves its correlation pattern; with a wrong phase choice, the transferred initialization is slower and less accurate than starting from random weights.","fun_headline_variants_meta":{"raw":{"variants":["Warm starts from small lattices beat random init","Blockwise tiling fastest for Ising ground states","Transfer learning scales neural quantum states","Reused weights cut time to larger ground states","Phase-aware tiling accelerates NQS training"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000802,"raw_usage":{"total_tokens":3534,"prompt_tokens":961,"completion_tokens":2573,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":2504}},"tokens_in":577,"tokens_out":2573,"duration_ms":22454,"temperature":1.0,"reasoning_tokens":2504,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:59:32.509074+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On the one-dimensional transverse-field Ising chain with $J_I = -2$ (antiferromagnetic phase), scale from 4 to 128 spins using the $(1,2)$-tiling, $(L,2)$-tiling, and a cold start with the paper's stopping criterion; the central claim fails if the $(L,2)$-tiling hot start does not reach the stopping criterion in less time and with lower mean relative energy error than the cold start over the 20 realizations.","supporting_citations":[{"cited_title":"NetKet: A Machine Learning Toolkit for Many-Body Quantum Systems","cited_arxiv_id":"1904.00031","evidence_quote":"It supplies the transfer-learning framing of reusing a trained model for a related task."},{"cited_title":"Abadi, P","cited_arxiv_id":null,"evidence_quote":"It supplies evidence that features transfer between same-architecture neural networks, motivating the hot-start initialization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It is the specific CPU implementation used as the timing baseline in the speed comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the GPU machine-learning platform on which the implementation runs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the quantum Monte Carlo reference energies and correlators for the two-dimensional model."}],"review_version":1}