{"id":"95ee0142-ef7f-4c6c-9063-a32d42e89ccd","arxiv_id":"2601.04402","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Penalty weights in a QUBO encoding act as thermodynamic control knobs, changing both solver success and irreversibility on a quantum annealer.","lead":"This paper shows that the penalty weights used to turn a scheduling problem into a QUBO optimization also change how much entropy and energy the quantum annealer dissipates. It suggests that encoding choices should be made with thermodynamics in mind, not just solution quality.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The thermodynamic inference pipeline (pseudo-likelihood β2 from non-equilibrium outputs, TUR bounds, and efficiency as a ratio of lower bounds) is unvalidated; without it the central claim that penalties control dissipation is unsupported.","rationale":"The reader's conditional verdict is appropriate. The paper has real independent support: the regime boundary in classical solvers, master-equation solution probabilities, and hardware reverse-annealing maps is consistent, and the asymmetry between p_sum and p_pair is structurally explained. The computational half of the claim — penalties determine feasibility and solver success — is reasonably supported. The thermodynamic half is where the argument is weakest. Every plotted entropy/work/heat/efficiency value is a TUR lower bound using a pseudo-likelihood β2 estimated from the same non-equilibrium outputs used to compute ΔE1. The paper explicitly concedes the outputs are generally non-equilibrium freeze-out samples, so the Gibbs assumption behind Eq. (23) is known to be violated. Since β2 enters multiplicatively in Eqs. (20)-(21), a biased β2 can rescale and reorder the bounds. Additionally, Eq. (22) is not a valid efficiency bound: the ratio of a lower bound on W and a lower bound on -Q does not bound -W/Q. The most direct settlement is to run the authors' own master-equation simulation through the same inference pipeline and compare TUR-inferred bounds to the exact thermodynamic quantities computed via Eqs. (26)-(33). If the pipeline does not reproduce the exact ordering, the experimental 'efficiency' maps are artifacts. This is a check the authors could have included and should be a condition for accepting the thermodynamic central claim.","tokens_in":20835,"tokens_out":10470,"duration_ms":113355,"concrete_test":"Use the authors' own adiabatic master-equation simulations (same JSP instance, β=10, τ=10 ns, same reverse schedules) to compute exact ⟨Σ⟩, ⟨W⟩, ⟨Q⟩ from the time-step decomposition in Eqs. (26)-(33). Then apply the pseudo-likelihood β2 estimator (23)-(24) and the TUR formulas (19)-(21) to the simulated output samples. Check whether the inferred efficiency ordering across (psum, ppair) reproduces the exact thermodynamic ordering. Separately, on the D-Wave data, re-estimate β2 from independent equilibrium-like samples at s=1 (long forward anneal or explicit thermalization) rather than from the same reverse-anneal outputs. If the TUR pipeline fails to match exact values, or if the thermodynamic regime boundaries shift with the β2 estimator, the central claim is an artifact of the inference method.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that QUBO penalties act as thermodynamic control knobs depends entirely on the inference in Sec. IV.A: the multivariate fluctuation theorem (13), the TUR bound (19), and the derived work/heat bounds (20)-(21). All require that the reverse-annealing statistics be governed by a factorized initial state and a single bath inverse temperature β2. But β2 is not measured; it is estimated by pseudo-likelihood (23)-(24) from the device's own final samples. The paper acknowledges (Sec. V.A) that these outputs are non-equilibrium freeze-out samples, not Gibbs states. A biased β2 changes the numerical value of every plotted bound through Eqs. (20)-(21), and can create or destroy the apparent efficiency ordering that is the main experimental result. There is also a technical inconsistency: Eq. (19) uses √⟨ΔE1²⟩ while Eqs. (20)-(21) use √var(ΔE1); and Eq. (22) forms the efficiency as a ratio of two lower bounds, which is not a valid bound on −W/Q. Before the thermodynamic conclusion is accepted, the inference pipeline needs validation with ground truth.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how QUBO penalty weights (p_sum for one-hot/sum constraints and p_pair for precedence constraints) affect both the computational hardness and the thermodynamic cost of solving a Job Shop Scheduling problem on a quantum annealer. On a D-Wave Advantage processor, the authors perform cyclic reverse-annealing experiments initialized from thermal samples, measure the stochastic processor energy change ΔE1, and apply a thermodynamic uncertainty relation (TUR) to derive lower bounds on entropy production, work, and heat. These bound maps are swept over a two-dimensional penalty plane for three reverse-annealing depths, and are complemented by classical solver sweeps and adiabatic master-equation simulations. The central claim is that the same encoding transitions that govern solver success also reorganize dissipation: weak penalties create low-energy infeasible manifolds, over-strong penalties suppress the effective problem energy scale and increase irreversibility, and the apparent thermodynamic efficiency is reduced in the computationally hard regimes. The paper concludes that QUBO penalties should be viewed as thermodynamic control knobs.","tokens_in":21132,"tokens_out":5639,"duration_ms":58874,"significance":"If the thermodynamic inference were validated, the paper would make a timely and useful contribution: it connects encoding choices in QUBO problems to the open-system thermodynamics of commercial quantum annealers, and it proposes an experimentally accessible diagnostic based on energy-change statistics. The combination of hardware experiments, classical solver benchmarks, master-equation simulations, and an open code repository is a definite strength. The qualitative observation that penalty strength reshapes both solution statistics and dissipation is plausible and corroborated by the simulations. However, the experimental thermodynamic conclusions rest entirely on TUR lower bounds with a fitted effective bath temperature β2, and the reported efficiency is formed from a ratio of lower bounds. These load-bearing aspects are not validated against ground truth, and the central experimental maps lack uncertainty quantification. The result is therefore not yet established at the level claimed.","major_comments":[{"comment":"The thermodynamic inference pipeline is not validated. The bounds in Eqs. (19)-(21) require the bath inverse temperature β2, which is estimated by pseudo-likelihood from the device's own final samples (Eqs. (23)-(24)). Section V.A explicitly acknowledges that these samples are non-equilibrium freeze-out outputs, not Gibbs states. Since the bounds scale with 1/β2 times g(...), a biased β2 can create or destroy the apparent ordering in Figs. 4-6. The authors should validate the pipeline in the master-equation simulations of Sec. IV.B, where β2 and the environment temperature are known: compute the TUR bounds from simulated ΔE1 moments and compare them with the directly integrated ⟨Σ⟩, ⟨W⟩, and ⟨Q⟩, and also test whether the pseudo-likelihood estimator recovers the true β2 under non-equilibrium sampling. Without such a ground-truth check, the central claim that penalties 'control' dissipati","section":"Sec. IV.A, Eqs. (19)-(22)"},{"comment":"There is an internal inconsistency in the TUR application. Eq. (19) uses √⟨ΔE1²⟩ in the argument of g, while Eqs. (20)-(21) use √var(ΔE1). For a nonzero mean, these quantities are different and only one expression can be the correct TUR. This discrepancy changes every plotted bound in Figs. 4-6. Please identify the correct form, justify it from the cited TUR, and recompute the results consistently.","section":"Sec. IV.A, Eq. (19) vs Eqs. (20)-(21)"},{"comment":"The efficiency is formed as a ratio of two lower bounds: Eqs. (20)-(21) give lower bounds on -⟨Q⟩ and ⟨W⟩, not estimates of these quantities. A ratio of lower bounds is not a valid bound on -⟨W⟩/⟨Q⟩, so the relative efficiency ordering in panel (e) is not established. If the intended claim is only that the two bounds move in the same direction, say so and do not call it thermodynamic efficiency; otherwise derive a rigorous bound on the ratio or report direct simulation values.","section":"Sec. IV.A, Eq. (22) and Sec. V.A, Fig. 4-6(e)"},{"comment":"The experimental maps are presented without any error bars or confidence intervals. The text interprets fine spatial 'speckling' as hardware sensitivity, but with no uncertainty quantification this could be sampling noise, especially near transitions where var(ΔE1) grows. Please report standard errors (e.g., bootstrapped) for at least ⟨ΔE1⟩, the TUR bound, and the efficiency, and demonstrate that the p_sum-dominated regime boundary is statistically significant rather than an artifact of finite sampling.","section":"Sec. V.A, Figs. 4-6"}],"minor_comments":[{"comment":"Typo: 'efficency' should be 'efficiency'. Also, Eq. (23) is typeset awkwardly with the equation number appearing inline ('Λ(β) =(23)'); it should be a normal numbered equation.","section":"Sec. IV.A"},{"comment":"References [49] and [53] both cite 'D-Wave samplers (2025)' with the same URL. These should be distinct entries or merged, with a clear description of what was accessed.","section":"References"},{"comment":"The simulation results use τ = 10 ns and β = 10 for 4-qubit instances, while the hardware experiments use τ = 10 μs and β1 = 10 on a 10-qubit instance. The captions and main text should state this difference explicitly and explain why the comparison is expected to be qualitative.","section":"Sec. IV.B, Figs. 7-8"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the unvalidated TUR/pseudo-likelihood inference pipeline. I would not accept the paper without a ground-truth validation in the simulations and a corrected, rigorously justified efficiency statement. The paper's core qualitative idea is plausible and within scope, and the simulations already provide the ingredients for the needed validation. The lack of error bars in the central experimental figures is also a barrier for a quantitative journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper maps the QUBO penalty parameter plane for a small Job Shop instance onto thermodynamic observables measured on a D-Wave device, using thermodynamic uncertainty relations to bound entropy production, work, and heat from energy-change statistics. That combination—systematic penalty-space sweep plus TUR-based inference on real hardware—is new, and the paper is honest about what is a bound and what isn't. The qualitative picture is plausible: weak penalties create low-energy infeasible states, strong penalties shrink the effective problem energy scale, and the encoding transition visible in solver success also shows up in the inferred dissipation. The p_sum vs p_pair asymmetry is explained structurally, and the same regime boundary appears across classical solvers and master equation simulations, which is the strongest evidence.\n\nThe soft spots are real but proportionate. β2, the effective bath inverse temperature, is estimated from the device's own final samples via pseudo-likelihood, and those samples are non-equilibrium freeze-out states. The paper acknowledges this, but a biased β2 propagates through every plotted bound and can create or destroy the efficiency ordering that is the headline result. The efficiency is also computed as a ratio of two lower bounds, which is not a valid bound on −W/Q; the paper wisely avoids making much of absolute values, but the relative ordering is still not bulletproof. There is also a technical inconsistency: Eq. (19) uses √⟨ΔE1²⟩ while Eqs. (20)–(21) use √var(ΔE1). Probably a typo, but it needs fixing. No error bars on the hardware maps makes it hard to tell how much of the \"speckling\" is signal. And the hardware runs are a single 10-qubit instance; the simulations use a different 4-qubit problem and forward annealing, so the corroboration is qualitative.\n\nThat said, the central claim—that penalty weights act as thermodynamic control knobs—is plausible, and the paper does not overstate it. It is an empirical study with clear limits, and the authors flag most of them. The right referee can push on the β2 validation and the bound-ratio issue. This deserves peer review, not a desk reject, and I would cite it for the encoding-to-dissipation connection once the inference concerns are addressed.","headline":"Useful empirical mapping of QUBO penalties to TUR-based thermodynamic diagnostics on D-Wave, plausible but the inference pipeline needs validation before the dissipation claims are fully trusted.","tokens_in":21605,"tokens_out":2200,"would_cite":true,"duration_ms":23582,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that QUBO penalty weights are thermodynamic control knobs: they set not only whether a quantum annealer finds feasible solutions but also how much entropy production, work, and heat the annealing cycle costs.","keywords":["quantum annealing","QUBO encoding","Job Shop scheduling","thermodynamic uncertainty relation","entropy production","reverse annealing","penalty weights","dissipation"],"falsifier":"Take the same Job Shop instance and reverse-annealing protocol, but record the joint distribution of processor and bath energy changes using an independent calorimetric or weak-measurement probe; if the measured ΔE2 violates the heat bound derived from the TUR, or if the pseudo-likelihood estimate of the bath temperature disagrees with the true bath temperature enough to change the sign of the inferred work/heat bounds near the feasible/infeasible boundary, the central thermodynamic reading is falsified.","tokens_in":20758,"feed_emoji":"⚛️","tokens_out":4277,"duration_ms":43987,"temperature":0.7,"pith_summary":"The authors ask whether the way a constrained optimization problem is translated into QUBO form affects not just solution quality but also the built-in dissipation of the annealer itself. Using a Job Shop scheduling instance with two tunable penalty weights, they find sharp transitions in feasibility and solver success, and show that the same transitions appear in inferred entropy production, work, and heat during reverse annealing. Weak penalties leave low-energy infeasible states that dominate sampling; excessive penalties shrink the effective energy scale and make the process more irreversible. The paper concludes that penalty tuning is fundamentally an energy-scale matching problem, and that sum-type constraints carry most of the thermodynamic leverage.","feed_headline":"Penalty weights set both annealer accuracy and heat cost","feed_subtitle":"Weak penalties conceal infeasible states; huge penalties shrink the energy scale and raise irreversibility.","key_machinery":"The load-bearing tool is the thermodynamic uncertainty relation (TUR): for a process with a fluctuation symmetry, the mean entropy production is bounded below by a function of the ratio of the mean to the standard deviation of any current, here the stochastic processor energy change ΔE1. Measuring only the first two moments of ΔE1 during a cyclic reverse-anneal protocol yields lower bounds on entropy production, work, and heat. The unknown bath temperature is estimated from the device's own output samples using pseudo-likelihood fitting of an Ising model, and the resulting bounds are checked against adiabatic master-equation simulations.","core_discovery":"The paper's central discovery is that the penalty weights used to translate a constrained scheduling problem into QUBO form do not merely gate whether an annealer returns feasible answers; they reshape the low-energy spectrum and, through it, the thermodynamic cost of the machine's operation. Sweeping the one-hot penalty p_sum and the precedence penalty p_pair, the authors find sharp boundaries between feasible and infeasible QUBO ground states, and show that these same boundaries appear as coordinated changes in inferred entropy production, work, and heat bounds measured in cyclic reverse-annealing runs on a quantum annealer. Weak penalties leave low-energy infeasible manifolds that dominat","pith_inferences":["If this correlation holds beyond Job Shop problems, the variance of ΔE1 in a short reverse-anneal scan could serve as a cheap pre-screening metric for encoding quality, before any full optimization campaign.","The paper's effective-temperature interpretation suggests a concrete design rule: choose penalty magnitudes so that the smallest meaningful energy gap in the encoded spectrum stays above the analogue control-error and thermal-noise scale of the specific device.","A natural extension is to combine the thermodynamic-efficiency ordering with time-to-solution or hybrid post-processing costs; if those orderings conflict, resource-aware encoding selection would need a multi-objective version of the claimed trade-off.","Because the TUR bounds are only bounds, tighter single-shot work measurements or calorimetric access to the bath would be needed to confirm whether the relative efficiency ordering survives at the level of true values, not just bounds."],"forward_implications":["The same quantities that mark computational hardness—the feasible/infeasible and split/unsplit transitions—also reorganize dissipation, so hardness and inefficiency are linked in a measurable way.","Tuning penalties is not an algorithmic detail; it changes the effective energy scale seen by the hardware, so 'as large as possible' is not a valid heuristic.","Because p_sum acts on the diagonal of the QUBO, encoding families should be expected to have asymmetric thermodynamic sensitivity; sum-type constraints are the dominant control.","Reverse-annealing energy statistics give a practical, device-agnostic route to building a thermodynamic phase diagram of a QUBO encoding space without access to bath observables."],"fun_headline_variants":["QUBO penalties tune both answer success and heat cost","Penalty weights in QUBO govern accuracy and entropy production","Annealing heat leaks tied to QUBO penalty strengths","Feasibility and irreversibility share the same QUBO knobs"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The reported entropy, work, and heat values all inherit the assumption that the annealer's output can be read through a multivariate fluctuation theorem with a single effective bath temperature estimated by pseudo-likelihood from the device's own samples; if the final samples are not approximately Gibbsian, or the system is not weakly coupled with a factorized initial state, the inferred thermodynamic quantities lose their physical meaning.","fun_headline_variants_meta":{"raw":{"variants":["QUBO penalties tune both answer success and heat cost","Penalty weights in QUBO govern accuracy and entropy production","Annealing heat leaks tied to QUBO penalty strengths","Feasibility and irreversibility share the same QUBO knobs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000238,"raw_usage":{"total_tokens":1360,"prompt_tokens":766,"completion_tokens":594,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":534}},"tokens_in":510,"tokens_out":594,"duration_ms":5641,"temperature":1.0,"reasoning_tokens":534,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T12:02:21.308441+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same Job Shop instance and reverse-annealing protocol, but record the joint distribution of processor and bath energy changes using an independent calorimetric or weak-measurement probe; if the measured ΔE2 violates the heat bound derived from the TUR, or if the pseudo-likelihood estimate of the bath temperature disagrees with the true bath temperature enough to change the sign of the inferred work/heat bounds near the feasible/infeasible boundary, the central thermodynamic reading is falsified.","supporting_citations":[],"review_version":1}