{"id":"9d81c17d-fa41-4697-a45a-3b5228f55cc0","arxiv_id":"2607.03208","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"FM+QA with data-driven Tchebycheff scalarization finds non-convex Pareto alloy mixtures on D-Wave hardware and matches classical FM+SA up to ~25 binary qubits, while one-hot encoding degrades earlier.","lead":"Researchers tested a factorization-machine plus quantum-annealing loop for multi-objective alloy recycling design, using Tchebycheff scalarization to find non-convex Pareto fronts of yield strength and thermal conductivity. The work maps when binary encoding and current D-Wave hardware remain competitive with classical simulated annealing for scrap-mixture optimization.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged model-proxy limitation.","rationale":"The reader's strongest claim accurately restates the paper's contribution and the weakest assumption correctly isolates the model-proxy gap that the authors themselves flag. Because the comparative computational evidence (binary vs one-hot, QA vs SA parity to 25 qubits, DDTS vs weighted-sum) is internally consistent and does not rest on unstated mathematical assumptions, no further load-bearing concern arises that would move the verdict. The recommended concrete test isolates the algorithmic claim from the material-model limitation without requiring new hardware or experimental campaigns. Therefore the existing CONDITIONAL verdict with high confidence remains appropriate.","tokens_in":19963,"tokens_out":472,"duration_ms":5518,"concrete_test":"Re-run the 20-qubit MOO case of Sec. 2.2 for 500 iterations with the same DDTS schedule but replace Thermo-Calc feedback by a fixed synthetic non-convex front whose ground-truth Pareto set is known analytically; if the recovered hypervolume and non-convex coverage remain statistically indistinguishable from the FM+SA baseline, the algorithmic claim is independent of the material-model fidelity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an empirical assessment that FM+QA+DDTS extends QUBO active learning to non-convex multi-objective alloy-mixture design and, under binary encoding, matches FM+SA on D-Wave Advantage up to 25 logical qubits while recovering non-convex front regions that weighted-sum scalarization misses. That claim is supported by repeated SOO runs across scales (Figs. 2–4), a 500-iteration MOO demonstration with hypervolume tracking against five FM+SA replicates (Fig. 5), and a brute-force grid check confirming quaternary mixtures occupy the non-convex segment (Fig. 6). The paper itself states that Thermo-Calc YS/TC estimates omit morphology, interconnectivity, impurities and full processing history (Introduction, Sec. 4.2) and that experimental validation lies outside scope; the reader already treats this as the weakest assumption and conditions the verdict on it. No additional load-bearing internal inconsistency, circular derivation, or unsupported comparative claim is present in the reported computational results.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript assesses an FM+QA active-learning workflow combined with data-driven Tchebycheff scalarization (DDTS) for multi-objective Pareto optimization of Al-alloy scrap mixtures, maximizing model-based yield strength and thermal conductivity. Across SOO benchmarks of increasing scale (3–5 mixing alloys, binary vs one-hot encoding, 9–25 logical qubits) it reports that binary encoding is markedly more efficient, that FM+QA on D-Wave Advantage matches FM+SA under matched settings, and that DDTS recovers non-convex front regions missed by weighted-sum scalarization; a 500-iteration MOO demonstration and a brute-force grid check corroborate the latter. Scaling, embedding overhead, TTS and near-term practical utility of QA are discussed critically.","tokens_in":20217,"tokens_out":945,"duration_ms":21456,"significance":"If the empirical findings hold, the work usefully extends QUBO-based active learning from single-objective to non-convex multi-objective materials-design problems and supplies concrete hardware benchmarks for a recycling-relevant use case. Strengths include 15-fold SOO repeats with mean/std, hypervolume tracking against FM+SA replicates, open-source code, explicit encoding and embedding details, and an honest appraisal of when current QA hardware remains competitive. These elements give the materials and quantum-optimization communities practical guidance even while true quantum advantage remains prospective.","major_comments":[{"comment":"Sec. 2.2 and Fig. 5d: the claim of comparable FM+QA vs FM+SA Pareto performance rests on a single FM+QA trajectory whose hypervolume lies inside the min–max envelope of five FM+SA runs. Given QA hardware stochasticity (embedding, noise, spin-reversal) and the 15-fold statistics used for SOO, additional independent FM+QA replicates are needed to place the MOO equivalence on the same footing; without them the central multi-objective claim remains only suggestive.","section":"Sec. 2.2 / Fig. 5d"},{"comment":"Sec. 2.1.4, Eq. (1) and Fig. 4b: TTS is computed under fixed, non-optimized SA schedules and workflow-specific QA anneal counts; the paper correctly notes that the observed QA advantage for N≤16 is therefore not a general solver benchmark. The abstract and conclusions nevertheless advertise a “critical perspective” on sizes at which QA “may become practically beneficial.” Either a limited SA annealing-time sweep on representative QUBOs or a sharper statement that no asymptotic crossover is claimed is required so that the TTS discussion does not over-reach the data.","section":"Sec. 2.1.4 / Eq. (1)"}],"minor_comments":[{"comment":"Table 1 lists 3.1 % resolution for model L while Sec. 2.1.3 writes “3.2 %”; reconcile the numbers.","section":"Table 1 / Sec. 2.1.3"},{"comment":"Fig. 2 caption swaps “cold/warm colors” relative to the text description of binary vs one-hot; correct for consistency.","section":"Fig. 2"},{"comment":"The local-search spin-flip mechanism is introduced only for one-hot runs (Sec. 2.1.2) yet is later applied also to binary L (Sec. 2.1.3); state the policy uniformly.","section":"Sec. 2.1.2–2.1.3"},{"comment":"Minor typographical inconsistencies appear (e.g., “F our alloy mixtures”, “3D spin glass”, missing spaces around units); a careful proof-read would remove them.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a solid methodological demonstration with open code and appropriate self-citation to the authors’ prior DDTS work. The Thermo-Calc proxy limitation is already flagged by the authors and does not undermine the computational claims; experimental validation is correctly scoped out. Fit for a materials-informatics or quantum-for-materials venue is good; the two major points above are readily addressable and do not require new hardware campaigns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is the first hardware run of DDTS-enabled FM+QA for non-convex multi-objective alloy-mixture design, plus a clean head-to-head of binary vs one-hot encoding and TTS on D-Wave Advantage. That is the real increment over Kitai-style single-objective FM+QA and over the authors’ own prior DDTS paper.\n\nWhat they do well: SOO is repeated 15 times with means and stds across five problem scales (9–25 binary qubits); MOO hypervolume is tracked against five FM+SA replicates; a brute-force 5% grid confirms that quaternary mixtures occupy the non-convex segment that weighted-sum misses. Encoding, embedding sizes, anneal counts, and spin-reversal settings are reported. Binary encoding clearly wins; one-hot degrades between 45–60 logical qubits as expected from embedding and feasibility constraints. They do not claim asymptotic quantum speedup and they state that SA was not exhaustively tuned for TTS. Code is on GitHub.\n\nSoft spots are real but proportional. Properties come from Thermo-Calc configured for rapid solidification; the paper itself notes that morphology, interconnectivity, impurities, and processing history are not fully captured, and experimental validation is out of scope. Raw Thermo-Calc data cannot be released for licensing reasons. Those are honest limitations, not hidden ones. Free parameters (anneal count, λ, FM rank, resolution) are standard for this workflow and are calibrated per sub-study.\n\nThis is for people who already care about QUBO active learning or scrap-based alloy design. It will not change industrial practice tomorrow, but it is a careful empirical assessment that a serious materials-informatics or quantum-optimization group should read. I would send it to referees; the computational claims are grounded enough to deserve that time. Engage if you work in this niche; otherwise skim the encoding and MOO figures.","headline":"Solid first hardware demo of DDTS-enabled FM+QA for non-convex multi-objective scrap-alloy design; binary encoding matches SA up to 25 qubits, with model-proxy limits already owned by the authors.","tokens_in":20866,"tokens_out":499,"would_cite":true,"duration_ms":5323,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Factorization machines plus quantum annealing can find non-convex Pareto-optimal scrap alloy mixes, matching classical annealing up to 25 qubits.","keywords":["QUBO-based optimization","multi-objective","Tchebycheff scalarization","factorization machine","alloy recycling","quantum annealing","Pareto front","binary encoding"],"falsifier":"Experimentally measure yield strength and thermal conductivity on a set of the reported Pareto-optimal scrap mixtures (especially the four-alloy non-convex points) under the same additive-manufacturing conditions; if those points are dominated by simpler binary or ternary mixes, or fail to appear on the experimental front, the computational Pareto claim fails.","tokens_in":20842,"feed_emoji":"♻️","tokens_out":1007,"duration_ms":9174,"temperature":0.7,"pith_summary":"This paper asks whether a factorization-machine surrogate turned into a QUBO and solved on a quantum annealer can handle multi-objective alloy recycling design, not just single-objective black-box problems. The authors combine the machine with data-driven Tchebycheff scalarization so that each active-learning step can target different trade-offs between yield strength and thermal conductivity, including non-convex parts of the front that weighted-sum scalarization cannot reach. They discretize mixing ratios of up to five commercial aluminum scraps, encode them with binary or one-hot bits, and compare D-Wave Advantage sampling against classical simulated annealing under matched active-learning loops. Binary encoding consistently works better; with it, quantum and classical samplers perform similarly up to 25 logical qubits and both recover well-resolved Pareto fronts that include complex four-alloy mixes. The practical message is that the workflow already functions as a competitive data-driven Pareto sampler on present hardware, while one-hot encodings and larger clique sizes remain limited by embedding noise and constraints.","feed_headline":"Quantum annealing finds non-convex scrap-alloy Pareto fronts","feed_subtitle":"FM+QA with Tchebycheff scalarization matches classical annealing up to 25 qubits and recovers complex multi-alloy mixes.","key_machinery":"Data-driven Tchebycheff scalarization (DDTS): a per-iteration preprocessing step that turns multi-objective property data into a single scalarized objective using randomized preference weights and a utopia point, so that the factorization machine produces a QUBO whose low-energy samples cover both convex and non-convex parts of the Pareto front.","core_discovery":"FM+QA combined with data-driven Tchebycheff scalarization extends QUBO-based active learning to non-convex multi-objective Pareto optimization of scrap-alloy mixtures. On the D-Wave Advantage system, binary-encoded instances up to 25 logical qubits match classical FM+SA performance and resolve non-convex front regions that weighted-sum scalarization misses, while one-hot encoding degrades earlier.","pith_inferences":["Hybrid pipelines that keep QA as a global sampler and hand the returned candidates to classical local search or constraint repair would likely push usable problem sizes past the current 25-qubit binary ceiling without waiting for denser hardware graphs.","The same DDTS-FM loop could be reused for other circular-materials problems (e.g., multi-scrap steel or battery-cathode blends) where competing properties are known to produce non-convex fronts.","Because each Thermo-Calc evaluation is expensive, the workflow’s value will rise sharply once experimental or higher-fidelity feedback replaces the model oracle, provided sample efficiency remains high."],"forward_implications":["QUBO-based active learning can now be applied to multi-objective materials design problems whose Pareto fronts contain non-convex regions.","Binary encoding of continuous mixture fractions is the practical default for FM+QA alloy problems; one-hot encodings become unattractive beyond roughly 45–60 logical qubits on present hardware.","For scrap-recycling design spaces of the size studied (up to five alloys, ~25 binary variables), quantum annealing is already a usable sampler inside the active-learning loop rather than a theoretical curiosity.","Expanding the same loop to more scrap streams or additional property objectives is limited mainly by sampler quality and calibration cost, not by the formulation itself."],"fun_headline_variants":["FM+QA recovers non-convex scrap-alloy Pareto fronts with Tchebycheff scalarization","Factorization machines plus quantum annealing tackle multi-objective alloy recycling","Binary-encoded FM+QA matches simulated annealing on 25-qubit non-convex Pareto fronts","Data-driven Tchebycheff enables QUBO optimization of competing scrap-alloy objectives","One-hot encoding falters earlier than binary in FM+QA alloy up-cycling searches"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The Thermo-Calc estimates of yield strength and thermal conductivity, tuned for rapid solidification and small grain size, are faithful enough proxies for the real scrap-based additive-manufacturing objectives that the discovered mixes remain meaningful once impurities, phase morphology, and processing history are present.","fun_headline_variants_meta":{"raw":{"variants":["FM+QA recovers non-convex scrap-alloy Pareto fronts with Tchebycheff scalarization","Factorization machines plus quantum annealing tackle multi-objective alloy recycling","Binary-encoded FM+QA matches simulated annealing on 25-qubit non-convex Pareto fronts","Data-driven Tchebycheff enables QUBO optimization of competing scrap-alloy objectives","One-hot encoding falters earlier than binary in FM+QA alloy up-cycling searches"]},"model":"grok-4.5","effort":"low","cost_usd":0.00447,"raw_usage":{"total_tokens":1293,"prompt_tokens":726,"num_sources_used":0,"completion_tokens":116,"cost_in_usd_ticks":44700000,"prompt_tokens_details":{"text_tokens":726,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":451,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":726,"tokens_out":116,"duration_ms":3966,"temperature":1.0,"reasoning_tokens":451,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T04:06:41.973543+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Experimentally measure yield strength and thermal conductivity on a set of the reported Pareto-optimal scrap mixtures (especially the four-alloy non-convex points) under the same additive-manufacturing conditions; if those points are dominated by simpler binary or ternary mixes, or fail to appear on the experimental front, the computational Pareto claim fails.","supporting_citations":[],"review_version":1}