{"id":"42deb873-08dd-4504-94c9-0ea5457895e5","arxiv_id":"2412.00504","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A quantum active learning workflow using quantum Gaussian process regression with projected and fidelity quantum kernels finds the global minimum of 4Al@Si11, but with no clear advantage over classical active learning and a search budget covering most of the isomer space.","lead":"This paper tests a quantum active learning loop in which quantum Gaussian process models suggest which doped nanoparticle structures to compute next, applied to finding the lowest-energy arrangement of 4 aluminum atoms on an 11-silicon cluster. It reports that the method finds the known global minimum structure, although classical active learning is often competitive or better.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Active-learning budget is nearly exhaustive (300 of 330 homotops) and no random baseline is included, so 'QAL found the GM in all 10 runs' is nearly forced and does not yet demonstrate feasibility.","rationale":"The reader's stated weakest assumption is descriptor quality (MBTR + PCA preserving energy-relevant structure). I agree that kernel-target alignment is unvalidated, but it is not the most load-bearing issue for the central claim. Because the budget exhausts 300 of the ~310 unexplored homotops, the GM will be found even by a random acquisition in ~97% of runs; after 10 runs the chance that random would also find it in all runs is ~72%. Therefore the observation 'GM found in all cases' cannot, by itself, support feasibility or efficiency. The missing random baseline makes the experiment insensitive to whether the quantum acquisition function is useful: a poor descriptor would still produce the same qualitative outcome. Adding a random-selection control and reporting first-hit distributions is the decisive check. This aligns with the reader's overall conditional verdict (their rationale also calls for a random control), so I do not change the verdict category, but I would emphasize the budget issue as the primary blocker rather than the descriptor embedding.","tokens_in":12771,"tokens_out":4952,"duration_ms":79380,"concrete_test":"Run the identical active-learning loop (same 10 random initial sets of 20 high-energy homotops, Ncycles = 60, Nselected = 5) but replace the acquisition function with uniform random selection without replacement from the unexplored homotops. Record the first cycle at which the putative GM enters the evaluated set for each run, and the average best-energy-vs-iterations curve. Compare this random baseline to the QAL/AL curves in Figs. 4 and 5. Also report the theoretical first-hit distribution: the GM's position among the ~310 unexplored homotops is uniform, so the median number of new calculations needed to find it is ~155, and 300 draws cover 96.8% of the space. If random reaches the GM at comparable or earlier budgets, or if the QAL curves are within run-to-run variance of random, the feasibility/efficiency claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Sec. 4 is that QAL is feasible because it found the putative GM of 4Al@Si11 in all 10 runs. However, the protocol in Sec. 3 uses Ncycles = 60 and Nselected = 5, i.e., 300 new DFTB local optimizations per run, on a fixed database of only 330 homotops. Each run starts with 20 random high-energy structures (E ≥ -12.2400 Ha), so the GM is almost certainly not in the initial set; the subsequent 300 selections are made from the remaining ~310 unexplored homotops. Even a uniform random rule selecting previously unvisited structures would encounter the GM with probability 300/310 ≈ 96.8% per run, and with probability ≈ 72% for all 10 runs. Thus the reported 'found in all cases' is almost inevitable irrespective of the acquisition function. The paper provides no random-selection control and no curve for the number of calculations required to first reach the GM; Figs. 4 and 5 only show average best-energy curves over the full 300-new-calculations budget. Without a baseline, the result is compatible with QAL being completely ineffective as a search guide (a near-exhaustive enumeration). The descriptor-quality concern raised by the reader is real, but it is secondary: even a blind acquisition succeeds under this budget, so the experiment cannot distinguish an informative from an uninformative kernel.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a quantum active learning (QAL) workflow for automatic structural determination of doped nanoparticles, combining quantum Gaussian process regression (QGPR) with fidelity and projected quantum kernels and two feature maps (YZ_CX and HighDim), implemented in the QMLMaterial software. The method is demonstrated on the 4Al@Si11 system, which has 330 DFTB-computed homotops, using 10 independent runs that each start from 20 random high-energy homotops and then perform 60 cycles of 5 new DFTB local optimizations (300 new calculations per run). The central claim, stated in Section 4, is that QAL found the putative global minimum in all 10 runs, demonstrating feasibility for structural determination of point-defect materials. The paper also compares QAL with classical active learning using Gaussian processes with two classical kernels.","tokens_in":13069,"tokens_out":2521,"duration_ms":27258,"significance":"If the claim were supported by a controlled comparison, the paper would be a useful demonstration of a quantum machine learning agent driving a materials discovery loop. The authors ship an implemented workflow (QMLMaterial) and use a concrete, reproducible-looking benchmark system with a known database of 330 homotops. The comparison of FQK and PQK quantum kernels, and of 4-qubit versus 8-qubit encodings, is a reasonable start toward understanding which quantum kernel designs help in active learning. However, the current experimental design does not provide evidence that the quantum acquisition function is actually guiding the search, because the budget is nearly exhaustive and no random-selection baseline is reported. As it stands, the paper is a proof-of-concept of the software plumbing rather than a demonstration of QAL as an effective search strategy.","major_comments":[{"comment":"The search budget makes the central claim nearly tautological. With Ncycles = 60 and Nselected = 5, each run performs 300 new DFTB local optimizations on a database of 330 homotops after removing the 20 initial structures. A uniform random rule that never revisits a structure would encounter the global minimum with probability about 300/310 ≈ 0.97 in a single run, and about 0.72 across all 10 runs. Therefore the observation that QAL found the GM in all 10 runs does not distinguish an informative acquisition function from blind enumeration. The paper needs a random-selection baseline and, preferably, a first-hit curve showing the number of new calculations required to reach the GM in each run; without these, the headline result is compatible with the acquisition function being completely ineffective.","section":"Section 3.1, Figs. 4 and 5"},{"comment":"The quantum kernel hyperparameters (including sigma = 0.0001 for PQK) and the feature-map/PCA choices were selected by grid search on the same fixed 4Al@Si11 energy database that is later used to evaluate QAL. This makes the reported MAE values and the eventual search performance partly fitted outcomes rather than independent predictions. The best-performance claims for QGPR-YZ_CX-PQK should be supported by a model-selection procedure that separates the data used for hyperparameter tuning from the data used for evaluation, or at least by a sensitivity analysis over the kernel hyperparameters.","section":"Table 1 and Section 2.5"},{"comment":"The paper does not validate whether the MBTR descriptor reduced by PCA to 4 or 8 components preserves the information needed for the quantum kernel to correlate with DFTB energies. No kernel-target alignment, distance-energy correlation, or other descriptor-quality check is reported before the descriptors are used to drive the active learning loop. Given that the random-budget argument shows the current experiment cannot detect a blind kernel, the paper should add a diagnostic that demonstrates the quantum kernel similarities carry information about the energy ordering, otherwise the QAL curves in Figs. 4 and 5 cannot be interpreted as evidence of learning.","section":"Sections 2.1 and 3.1"},{"comment":"The comparison between methods is made visually from average-energy curves without error bars or statistical significance tests. With 10 independent runs, the differences between some curves (for example, the QAL-PQK and QAL-FQK curves in Fig. 4) may be within run-to-run variability. The authors should report standard deviations or confidence intervals, and ideally a paired test across runs when comparing methods on the same initial conditions.","section":"Figures 4 and 5"}],"minor_comments":[{"comment":"The RBF kernel is introduced with the sentence 'The RBF kernel is given by Eq. 1' but the displayed formula is numbered Eq. (3); the cross-reference should be corrected.","section":"Section 2.2, Eq. (3)"},{"comment":"The projected quantum kernel formula contains an unfinished expression 'ρ_k(x_i) =' immediately before the text; the definition of the 1-RDM and the trace operator should be written out completely.","section":"Section 2.3, Eq. (5)"},{"comment":"The circuit name is written inconsistently as 'YC_ZX' in one place and 'YZ_CX' in others; also Fig. 3 is referenced in the text as 'Fig. X' and 'Fig. Y' in Section 2.5, and the figure captions should be checked.","section":"Section 3.1"},{"comment":"The text alternates between 'Tab. 1' and 'Tab. 2' for what appears to be the same hyperparameter table; the numbering and the associated explanations should be unified.","section":"Section 2.5 and Table 1"},{"comment":"There are several typographical issues, including 'Hatree' for Hartree, 'strutural' for structural, '4AL@Si11' for 4Al@Si11, and duplicated or malformed references in the reference list; a careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The core problem is that the experiment, as designed, cannot support the paper's central feasibility claim because the search budget is nearly exhaustive and there is no random baseline. This is not a matter of the quantum versus classical comparison being wrong; it is that the reported success metric is insensitive to the acquisition function. I would ask the authors to add a random-selection baseline, report first-discovery statistics, and either reduce the budget to a regime where the search is not trivially exhaustive or justify the current budget as deliberately matching the database size. The hyperparameter circularity is also a serious concern that should be addressed with an honest model-selection protocol. The paper may be suitable for a specialized quantum machine learning venue once these controls are in place."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The stress-test note is right: the paper's main claim — that QAL finds the putative GM in all 10 runs — is almost forced by the search budget. With 60 cycles and 5 selections per cycle, each run covers 300 of the 330 homotops, and starting with 20 random high-energy structures leaves about 310 unexplored. Even a uniform random rule that avoids repeats would find the GM with about 97% probability per run, and with about 72% across all 10 runs. Without a random-selection baseline, the result is compatible with the agent being no better than blind enumeration. The paper never reports the number of calculations needed to first hit the GM, only averaged best-energy curves over the full 300-new-calculations budget. That is a load-bearing gap, not a stylistic one.\n\nTo be fair, the paper does what it says: it implements QAL with two quantum kernels, two feature maps, and 4- and 8-qubit encodings, and compares with two classical GPR kernels. The method description is clear, and the authors are honest that no quantum advantage is claimed and that further work is needed. The hyperparameter grid search on the same fixed database is a real circularity, though it mainly affects which kernel/sigma look best, not the core feasibility claim. The missing descriptor validation (MBTR+PCA) is secondary: if the embedding were poor, the agent would be blind, but the budget would still save it, so the experiment cannot distinguish blind from informed.\n\nThe paper is an incremental case study, not a breakthrough. But it is a legitimate extension of the authors' own QAL program, and the comparison of kernel types is useful for practitioners. The fix is straightforward: add a random-selection and a fixed-order baseline, reduce the budget so the search is not exhaustive, and report first-hit curves and variance. Release the code and data if possible. With those additions, the feasibility claim could be supported; as written, it is too weak to stand alone.\n\nFor a reader working in quantum ML for materials, this is worth a quick read as an example of the workflow, but I would not cite it as evidence that QAL guides structural search. I'd send it to peer review, because the method and system are of interest and the weaknesses are addressable — but I would expect major revision before publication.","headline":"Near-exhaustive search budget makes the central feasibility claim nearly forced; the paper is an honest case study but needs a random baseline to support its conclusion.","tokens_in":13625,"tokens_out":2730,"would_cite":false,"duration_ms":27055,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a quantum active learning loop built on quantum Gaussian process regression guided DFTB structure searches to the known global minimum of 4Al@Si11 in all 10 independent runs.","keywords":["quantum active learning","quantum Gaussian process regression","doped nanoparticles","global minimum structure search","quantum kernels","MBTR descriptor","DFTB","structural determination"],"falsifier":"Compute a kernel-target alignment, such as centered kernel alignment, between the FQK and PQK kernel matrices on all 330 4Al@Si11 homotops and the vector of DFTB energies. Near-zero alignment would mean the kernel's notion of similarity carries essentially no energy information, implying the QAL loop's success came from its search dynamics rather than from the quantum kernel. A complementary check: rerun the same QAL loop with the energy labels randomly permuted; if the loop still appears to reach the global minimum under shuffled labels, it is not using the label information at all.","tokens_in":12601,"feed_emoji":"⚛️","tokens_out":11562,"duration_ms":97761,"temperature":0.7,"pith_summary":"Quantum active learning (QAL) is proposed as a way to find the most stable structure of a doped nanoparticle without exhaustively computing every possibility. The paper's agent is a quantum Gaussian process regression model that, at each cycle, chooses which unobserved dopant arrangements to compute next with density-functional tight-binding (DFTB), then retrains on the new energies. The test bed is 4Al@Si11, an 11-silicon cluster doped with four aluminum atoms, whose 330 arrangements (homotops) all have known DFTB energies, so every search can be scored against a known global minimum. Starting from 20 random structures with energies far above the global minimum, the QAL loop found the putative global minimum in all 10 independent runs. The claim is feasibility: a quantum kernel-based regression can steer structural search in a data-scarce setting, even where it does not beat classical baselines.","feed_headline":"Quantum agent lands on global minimum in all 10 structure searches","feed_subtitle":"The search agent steers DFTB calculations to the known lowest-energy structure of 4Al@Si11 in every run.","key_machinery":"The load-bearing object is the quantum active learning loop: a cycle in which a quantum Gaussian process regressor acts as a decision-making agent, selecting from the unexplored space of homotops the next candidates to be evaluated by DFTB local optimization, after which the observed energies are added to the training set and the regressor is refit. The regressor's similarity measures are quantum kernels — a fidelity quantum kernel $k_{FQK}(\\boldsymbol{x}_i,\\boldsymbol{x}_j) = \\langle \\varphi(\\boldsymbol{x}_i)|\\varphi(\\boldsymbol{x}_j)\\rangle$ and a projected quantum kernel $k_{PQK}(\\boldsymbol{x}_i,\\boldsymbol{x}_j) = \\exp(-\\gamma \\sum_{k,P} \\{ \\mathrm{tr}[P\\rho_k(\\boldsymbol{x}_i)] - \\mathrm{tr}[P\\rho_k(\\boldsymbol{x}_j)]\\}^2)$ built from single-qubit reduced density matrices — produced by two data-encoding circuits, YZ_CX and HighDim. Each homotop is described by the many-body tensor representation (MBTR), reduced by principal component analysis to 4 or 8 components, which fixes the number of qubits. The loop closes by updating the database and repeating until a preset number of cycles is reached; replacing QGPR by a classical Gaussian process regressor defines the classical active learning baseline.","core_discovery":"On the paper's own terms, the central discovery is that a quantum active learning loop built on quantum Gaussian process regression (QGPR) is feasible for automatic structural determination of point-defect materials. Concretely, QAL using the projected quantum kernel (PQK) with the YZ_CX feature map performed best among the quantum variants and found the putative global minimum of 4Al@Si11 in all 10 independent runs, even though each run started from 20 random homotops with energies at or above -12.2400 Hartree, deliberately far from the global minimum. With 4-qubit circuits the quantum and classical searches were competitive; with 8 qubits the classical Gaussian process with a dot-product kernel reached the global minimum fastest (in about 40 new calculations), yet the QAL variants still converged there. The authors note throughout that quantum kernel hyperparameters were kept fixed while classical kernel hyperparameters were automatically re-optimized as data accumulated, and they ascribe part of the classical advantage to that asymmetry. The whole demonstration is carried out in a noise-free quantum computing framework, so the claim is about the method's feasibility, not about hardware performance or quantum advantage.","pith_inferences":["The paper establishes feasibility, not advantage: its own data show the classical Gaussian process with a dot-product kernel and 8 principal components reached the global minimum in about 40 new calculations, faster than any quantum variant, so a reader should take the result as 'quantum active learning works' rather than 'quantum is better.'","The comparison is asymmetric in a way the authors flag: quantum kernel hyperparameters were fixed during the loop while classical kernel hyperparameters were re-optimized as data grew; re-running QAL with the quantum kernel's parameters tuned inside the loop is a direct test of whether the classical edge is inherent or an artifact of the fixed configuration.","A quantitative test that would separate descriptor quality from kernel expressivity is measuring the alignment between each quantum kernel and the DFTB energy labels on the full 330-homotop set; that analysis is absent from the paper and would explain why PQK outperformed FQK.","The sensible next application is a homotop space too large to enumerate fully (bigger clusters, more dopants, or vacancy sites), where the metric is total DFTB effort to reach an energy threshold; that is where a scarce-data search loop could show real value."],"forward_implications":["The paper's claim implies that an autonomous quantum-agent loop can replace exhaustive enumeration for small doped clusters: given a descriptor and an energy method, the loop reaches the known global minimum without visiting all 330 homotops.","QAL with the projected quantum kernel and YZ_CX feature map is presented as the best quantum variant, improving when the circuit grows from 4 to 8 qubits; if the claim holds, going to more qubits is a plausible route to better search performance.","The same QAL procedure is claimed to transfer to other doped nanoparticles and solids with point defects, since the loop is independent of the energy method (DFT or DFTB) and of the specific cluster.","Because all QAL variants found the global minimum even from deliberately poor starting populations, the method is presented as reliable across random initial data selection."],"supporting_citations":[{"why":"Supplies the 330 4Al@Si11 homotops with their DFTB energies, the benchmark database the QAL search is scored against.","marker":"[2]"},{"why":"The prior quantum active learning formulation (QSVR/QGPR with quantum kernels) that this work extends to automatic structural determination.","marker":"[20]"},{"why":"Quantum Gaussian process regression, the regression algorithm powering the QAL agent.","marker":"[23]"},{"why":"The sQUlearn quantum computing framework that provides the fidelity and projected quantum kernels and the YZ_CX and HighDim feature maps.","marker":"[9]"},{"why":"Introduces the projected quantum kernel used to define the PQK similarity measure in Eq. 5.","marker":"[8]"},{"why":"The many-body tensor representation (MBTR) descriptor used to encode each homotop's structure.","marker":"[29]"},{"why":"QMLMaterial, the software in which the QAL loop is implemented.","marker":"[26]"}],"fun_headline_variants":["Quantum active learning finds doped Si optimum in all 10 runs","QAL with quantum kernel hits global minimum 10/10 times","Quantum GP guides 10/10 searches to 4Al@Si11 minimum","Quantum active learning: perfect record on doped nanoparticle"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire search presupposes that, after MBTR encoding and PCA compression, structures that the quantum kernel judges similar actually have similar DFTB energies; if the compressed descriptors scramble the energy ordering, the quantum agent's selections are barely more informative than random picks, and the loop's success would not be attributable to the learning.","fun_headline_variants_meta":{"raw":{"variants":["Quantum active learning finds doped Si optimum in all 10 runs","QAL with quantum kernel hits global minimum 10/10 times","Quantum GP guides 10/10 searches to 4Al@Si11 minimum","Quantum active learning: perfect record on doped nanoparticle"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000664,"raw_usage":{"total_tokens":3079,"prompt_tokens":1038,"completion_tokens":2041,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":654,"completion_tokens_details":{"reasoning_tokens":1968}},"tokens_in":654,"tokens_out":2041,"duration_ms":14569,"temperature":1.0,"reasoning_tokens":1968,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:18:17.573580+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute a kernel-target alignment, such as centered kernel alignment, between the FQK and PQK kernel matrices on all 330 4Al@Si11 homotops and the vector of DFTB energies. Near-zero alignment would mean the kernel's notion of similarity carries essentially no energy information, implying the QAL loop's success came from its search dynamics rather than from the quantum kernel. A complementary check: rerun the same QAL loop with the energy labels randomly permuted; if the loop still appears to reach the global minimum under shuffled labels, it is not using the label information at all.","supporting_citations":[],"review_version":1}