{"id":"d1953acf-8bd8-4b3d-b850-5f9cc39fac98","arxiv_id":"1908.08968","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Violations of global passivity inequalities, measured on IBM hardware, detect an engineered qubit-environment coupling in systems of up to four qubits.","lead":"The authors ran thermodynamic inequality tests on IBM quantum processors and found that the tests can flag when a qubit environment is coupled to a multiqubit system. This offers a measurement-based way to detect non-unital noise without full state tomography.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed detection intervals omit systematic uncertainty from readout inversion and from the 0.005 L2 deviation of the initial state from a product thermal state; a unital or unitary evolution on a slightly non-thermal state could plausibly produce the observed negative Delta<F>.","rationale":"The reader's conditional verdict is well-founded. The unital/random-unitary misstatement in the main text is a genuine theoretical error, but the inequality Delta<F> >= 0 actually holds for all unital channels via majorization, so the contrapositive used for detection survives. The post-hoc choice of alpha = 3 and the delta-window weakens the blindness of the test but does not invalidate the demonstration, especially since the decoupled control yields positive values. The load-bearing gap is the systematic error budget: the test constructs the observable from the same noise-corrected data that are used to evaluate it, and the reported error bars include only sampling noise. The stated 0.005 L2 distance to a product thermal state is an indication of preparation error, not a proof that the inequality remains valid under that deviation. A bootstrap and perturbation analysis, as proposed, would either restore confidence or reveal that the detection window lies within systematic uncertainty. Therefore the CONDITIONAL verdict should stand pending this check.","tokens_in":22693,"tokens_out":19393,"duration_ms":208285,"concrete_test":"Perform a nonparametric bootstrap over the ten Melbourne batches: for each bootstrap sample, resample the raw counts, draw M from a binomial model of its column calibrations, re-invert to obtain {p_i^s} and {p'_i^s}, reconstruct B and F_{3,delta}, and recompute Delta F_{3,delta} for delta in [1.5, 2.8]. Also simulate the preparation circuit with a unitary or unital noise model plus a random initial-state perturbation of L2 size 0.005, including off-diagonal coherences, and check whether Delta F_{3,delta} can fall below the reported 3-sigma band. If the bootstrap confidence interval includes zero in the coupled case, or if the unitary simulation yields a negative interval, the detection is not robust to the stated systematic uncertainties.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In the Melbourne experiment, the observable B is defined via B_i = -ln(p_i^s), where {p_i^s} are the initial populations recovered by applying M^{-1} to the observed counts (Supplemental, 'Detector noise and characterization of initial states'). The same {p_i^s} are then inserted into Eq. (4) to evaluate Delta<F_{alpha,delta}>. The confidence intervals (Supplemental Eqs. S12-S15) treat F_{alpha,delta} as a fixed operator and include only shot-noise variances of the initial and final populations; they do not propagate uncertainty in M, statistical error of the 10-batch estimate of {p_i^s}, or the 0.005 L2 distance between the recovered populations and the nearest product thermal state. A biased M^{-1} (e.g., from crosstalk or drift) or small initial coherences or correlations would change the eigenvalues and ordering of B, and could make Delta<F_{3,delta}> negative even for a unital or unitary evolution. The decoupled control uses a different circuit (no environment CNOTs), so it does not isolate the additional gate errors from the engineered environment. Thus the central experimental claim, that a negative Delta<F> unambiguously signals a heat leak, lacks a demonstrated robustness bound to these systematic errors.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes to detect a hidden environment coupled to a quantum system by looking for violations of thermodynamic inequalities. For a product thermal initial state ρ_s = ⊗ e^{-β_i H_i}/Z_i, passive observables F that are diagonal in the energy basis with non-decreasing eigenvalues are predicted to satisfy Δ⟨F⟩ = Tr[F(ρ'_s − ρ_s)] ≥ 0 for unital evolutions. A negative Δ⟨F⟩ is therefore interpreted as a heat leak, i.e., a non-unital component. The authors implement tests with F_{α,δ} = (B − δI)^α and with a deformed observable B_def on IBM quantum processors: a four-qubit system plus one-qubit environment (Melbourne) and a three-qubit system plus one-qubit environment (Essex). They report detection windows (Melbourne 1.8 ≲ δ ≲ 2.6 for F_{3,δ}; Essex 0.8 ≲ δ ≲ 2.1 for a deformed F^{(def)}_{5,δ}) that are negative by more than three standard deviations when the environment is coupled, and non-negative in the decoupled-environment controls. A supplemental analysis uses classical shadows to argue that evaluating these tests requires only O(poly(n)) measurements for fixed α and δ.","tokens_in":22973,"tokens_out":16345,"duration_ms":175558,"significance":"If the claims hold, the work offers a scalable, tomography-free diagnostic for non-unital errors in quantum devices, with the unusual feature that the test is independent of the details of the noise model. Strengths of the manuscript include genuine hardware experiments on two IBM processors, decoupled-environment control runs, and a scaling argument that correctly reduces the estimation of (B − δI)^α to local observables whose shadow norm is independent of n. The central inequality itself is more robust than the paper's stated justification: it holds for all unital maps via the doubly stochastic action on diagonals, not only for mixtures of unitaries, so the detection principle is theoretically sound. However, the experimental evidence for unambiguous detection is not yet airtight because the readout-calibration uncertainty and the non-unital gate-error component of the environment CNOTs are not quantified. With a corrected theoretical statement and the requested robustness analysis, this would be a worthwhile contribution to quantum device characterization.","major_comments":[{"comment":"The statement 'For any unital evolution, the final state ρ′s can be written as ρ′s = ∑i qi U(i)s(ρs)U(i)†s' is false for dimensions d > 2, and the following 'i.e.' incorrectly equates 'non unital' with 'not representable as a mixture of unitaries'. In the present four-qubit setting (d = 16), this is a substantive mathematical error. The inequality Δ⟨F⟩ ≥ 0 does, however, hold for every unital map when ρs is diagonal in the eigenbasis of F, because the diagonal of the output state is a doubly stochastic image of the initial diagonal and F has non-decreasing eigenvalues. I recommend replacing the mixture-of-unitaries justification with this majorization argument and defining non-unitality in the standard sense; the detection logic (violation ⇒ non-unital) then remains valid.","section":"Main text, paragraph after Eq. (2)"},{"comment":"The error bars do not include the uncertainty in the readout calibration matrix M or the statistical error in the estimate of {p_i^s} from the ten batches, even though B and F are built from those same {p_i^s}. Equations (S12)–(S15) treat F_{α,δ} as a fixed operator and use only the shot noise of the initial and final populations. A biased M^{-1} (from crosstalk or drift) or the reported 0.005 L2 distance between the recovered populations and the nearest product thermal state could reorder the eigenvalues of B and make Δ⟨F_{3,δ}⟩ negative for a unitary or unital evolution. Please provide a quantitative bound: propagate the uncertainty in M (for example, by bootstrapping the calibration data) and verify that the detection intervals 1.8 ≲ δ ≲ 2.6 (Melbourne) and 0.8 ≲ δ ≲ 2.1 (Essex) survive that propagation.","section":"Supplemental 'Detector noise and characterization of initial states' and 'Statistical error'; main text around Eq. (4)"},{"comment":"The decoupled-environment control removes the environment CNOTs, so it does not control for the additional gate errors introduced by those CNOTs in the coupled circuit. The observed negative Δ⟨F⟩ could therefore be caused by the non-unital component of the CNOT gate errors themselves rather than by the engineered thermal coupling. This does not invalidate the method—any actual non-unital component is a heat leak—but it weakens the specific physical claim that the source is the intended environment. An additional control that varies the coupling strength (the number of environment CNOTs or the environment temperature) and shows a systematic dependence of the detection interval would substantially strengthen the demonstration.","section":"Main text, Figs. 2 and 3 and the discussion of the decoupled controls"},{"comment":"The derivation gives Δ⟨F⟩ = −Σ_j Δf_j arξ_j, so a negative Δ⟨F⟩ requires a positive arξ_j. Two sentences later the text states that Δ⟨F⟩ < 0 'only if arξ_j < 0', which is inconsistent with the displayed formula. Since the proof of Theorem 1 is consistent with the formula, this appears to be a sign typo, but it should be corrected because it obscures the logic of the sensitivity claim.","section":"Supplemental, section on sensitivity of heat leak tests, after Eq. (S7)"}],"minor_comments":[{"comment":"In the sentence about the Essex processor, the text says 'n = nMel = 4 batches'; the symbol should be nEss (or a similar distinct label) to avoid confusion with the Melbourne batch count.","section":"Supplemental, 'Statistical error'"},{"comment":"There are several typographical errors: 'enviroment' in the abstract, 'in constrast' in the introduction, and 'not a applied' in the supplemental material. A careful proofreading pass is needed.","section":"Abstract and main text"},{"comment":"The sentence 'for 0 ≤ α ≤ 4 no detection occurs with the observables B^α' should be qualified as 'no unambiguous detection within three standard deviations', since the supplemental data show that the mean value of Δ⟨B^α⟩ becomes negative for α ≳ 3.4 even though the confidence interval still contains positive values.","section":"Main text, paragraph after Fig. 2"},{"comment":"The expansion in Eq. (6) is written with a sum over k from 0 to α, and the sign discussion is helpful; it would be even clearer to state explicitly that the bound on δ from the passivity condition must be imposed when α is even, as is later done for F_{2,δ}.","section":"Main text, Eq. (6)"}],"recommendation":"major_revision","confidential_remarks":"The false claim that every unital evolution is a convex mixture of unitaries is the kind of statement that a knowledgeable referee will flag immediately; it should be fixed before publication. The main reason I do not recommend minor revision is the missing propagation of readout-calibration uncertainty, which is load-bearing for the claimed unambiguous detection. I see no reason to doubt the authors' good faith: the self-citation to Ref. [27] is appropriate because the inequalities originate there, and the experimental controls are a reasonable first step. The paper fits the scope of the journal and the central idea is worth publishing once the robustness analysis is supplied."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one-liner: this is the first experiment I know of that actually runs these passive-observable \"heat leak\" tests on superconducting hardware, and it shows the idea can work. But the paper oversells the word \"unambiguous\": the reported confidence intervals ignore exactly the systematics that could produce a false violation.\n\nWhat's genuinely new: the Melbourne and Essex experiments, the passivity-deformed observables (setting one beta to zero), and the supplemental scaling bound using shadow tomography. The authors also do the right thing in the controls: they run the environment decoupled, they simulate the ideal circuit, and they check that the measured initial state is close to a product thermal state. They are honest in the discussion that the tests miss some non-unital errors and don't touch dephasing. The polynomial measurement cost is correctly stated as conditional on fixed alpha and delta ranges.\n\nThe first soft spot is a real theoretical error: the text says every unital evolution is a convex mixture of unitaries. That's false for d>2. The inequality itself probably survives because for any unital channel the dephased final state is majorized by the initial product state, and monotone functions of B then give Delta<F> >=0. But the proof as written is wrong and should be fixed.\n\nThe bigger issue is experimental. In Melbourne, B is built from the initial populations recovered from M^{-1}, and the same populations are used to evaluate Delta<F>. The error bars in the supplemental are shot-noise only. They don't propagate uncertainty in the readout matrix, the 10-batch spread in the initial populations, or the 0.005 L2 distance from the nearest product thermal state. If M^{-1} is slightly biased, the function -ln p can change a lot for low-population states, and a unital or even unitary evolution on a slightly non-thermal state could produce a negative Delta<F> in the reported window. The decoupled control is not a clean isolation of the engineered environment: it removes the CNOTs, so it also removes their gate errors. So the central claim, that a negative Delta<F> unambiguously flags a heat leak, lacks a demonstrated robustness bound.\n\nThere is also some post-hoc observable selection: they scan alpha and delta and report the region that works, without a multiple-comparison correction. That's a minor issue given the decoupled control and simulation, but worth noting.\n\nWho should read this: people building quantum-device validation tools, and quantum thermodynamics experimentalists. It deserves a serious referee; the experiments are real and the concept is useful. My recommendation: revise, fix the unital statement, propagate the calibration and initial-state uncertainty, and then it's publishable.","headline":"A credible first experimental demonstration of thermodynamic heat-leak detection on IBM hardware, but the 'unambiguous' claim outruns the error bars because calibration and initial-state systematics are not propagated.","tokens_in":23515,"tokens_out":7075,"would_cite":false,"duration_ms":65272,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Thermodynamic inequalities can act as black-box detectors: a negative passive-observable test reveals that a quantum system has been touched by a hidden environment.","keywords":["thermodynamic inequalities","non-unital dynamics","heat leak detection","passive observables","global passivity","quantum error diagnostics","shadow estimation","quantum thermodynamics"],"falsifier":"Prepare a product thermal state on a processor, apply only unital channels (random mixtures of unitaries or depolarizing noise) with the ancilla environment decoupled, and check that every $\\Delta\\langle F_{\\alpha,\\delta}\\rangle$ stays nonnegative; then deliberately perturb the readout-correction matrix $M$ (for example, swap the detection probabilities of one qubit) and see whether a negative test appears. If a calibration error alone can produce a violation, the 'hidden-environment' reading of the inequality is not self-contained.","tokens_in":22455,"feed_emoji":"🌡️","tokens_out":8321,"duration_ms":77320,"temperature":0.7,"pith_summary":"The paper claims that a violation of a thermodynamic inequality can act as a black-box detector for a hidden environment: if the change in the mean value of a suitably passive observable is negative, the system's evolution cannot be unital, so it must have leaked heat. For initial states that are products of thermal states, the relevant observables are built from $B=-\\ln\\rho_s$, and experiments with four-qubit and three-qubit systems coupled to a single-qubit environment show that negative values appear only when the environment is switched on. The same data show that shifted and deformed passive observables can detect leaks that the standard Clausius inequality and simple powers of $B$ miss. The paper also proves that evaluating these tests requires a number of measurements that scales polynomially with system size, in contrast to tomography-based diagnostics.","feed_headline":"Thermodynamic violation exposes a hidden quantum environment","feed_subtitle":"Tests from one passive observable catch hidden heat leaks without tomography, at polynomial cost.","key_machinery":"The load-bearing object is the family of passive observables $F_{\\alpha,\\delta}=(B-\\delta I)^\\alpha$, where $B=-\\ln\\rho_s$ and $\\delta$ is a shift chosen so that the eigenvalues of $F$ remain a non-decreasing function of the eigenvalues of $B$. Passivity guarantees $\\Delta\\langle F_{\\alpha,\\delta}\\rangle\\ge0$ under any unital transformation, so a negative value signals non-unitality; the expansion $\\Delta\\langle(B-\\delta I)^\\alpha\\rangle=\\sum_{k=0}^\\alpha \\binom{\\alpha}{k}(-1)^k\\delta^k\\Delta\\langle B^{\\alpha-k}\\rangle$ shows why the shift works, since odd-$k$ terms carry negative signs and can outweigh positive even-$k$ contributions. A second mechanism, passivity deformation, replaces $B$ by $B_{\\rm def}=\\sum_i \\beta_i^{({\\rm def})}H_i$ with effective temperatures that keep the observable passive, which reweights the expansion and enables detection when unshifted observables fail. On the resource side, the paper uses shadow estimation of local $k$-qubit operators $\\tilde{H}_{k,\\{i_m\\}}$ to show that $\\Delta\\langle F_{\\alpha,\\delta}\\rangle$ can be evaluated with $N\\sim O\\!\\left(\\frac{\\log(n/\\mu)}{\\varepsilon^2}(n+\\delta)^{2\\alpha}\\right)$ measurements, polynomial in the number of qubits.","core_discovery":"The central claim is that $\\Delta\\langle F\\rangle=\\mathrm{Tr}[F(\\rho_s'-\\rho_s)]<0$ for a passive observable $F$ is a certificate of non-unital dynamics, which the authors call a heat leak. Given an initial product of thermal states $\\rho_s=\\otimes_i e^{-\\beta_i H_i}/\\mathrm{Tr}(e^{-\\beta_i H_i})$, the observable $B=-\\ln\\rho_s$ satisfies $\\Delta\\langle B\\rangle\\ge0$ for every unital evolution; the same holds for every $F$ that commutes with $B$ and has eigenvalues given by a non-decreasing function of the eigenvalues of $B$. The experiments implement engineered environment couplings on small superconducting processors: with the environment coupled, the test $\\Delta\\langle F_{3,\\delta}\\rangle$ with $F_{3,\\delta}=(B-\\delta I)^3$ becomes negative in the interval $1.8\\lesssim\\delta\\lesssim2.6$, while $\\Delta\\langle B\\rangle$ and $\\Delta\\langle F_{2,\\delta}\\rangle$ do not detect it; in a second experiment, a deformed observable $B_{\\rm def}=\\beta_0(H_1+H_3)$ makes $\\Delta\\langle F^{({\\rm def})}_{5,\\delta}\\rangle$ negative for $0.8\\lesssim\\delta\\lesssim2.1$. A theoretical analysis shows that all these tests can be estimated from mean values of local observables, with a measurement count scaling polynomially in the number of qubits.","pith_inferences":["The same violation-as-diagnostic principle should transfer to other platforms—bosonic modes, trapped ions, or photonic circuits—where $B=-\\ln\\rho_s$ can be defined and passive observables measured, although the paper only demonstrates qubit processors.","Because the paper does not propagate uncertainty in the readout-correction matrix $M$ into the reported confidence intervals, a natural extension is a sensitivity analysis quantifying how much calibration bias can fake or mask a heat leak.","Sweeping $\\alpha$ and $\\delta$ yields a curve of test values that may act as a coarse fingerprint of the non-unital part of a noise channel, going beyond binary detection toward classification.","The same thermodynamic logic could be applied to detect dephasing by choosing passive observables in rotated bases, since pure dephasing is unital and will not trigger the current tests; the paper lists dephasing as an open problem."],"forward_implications":["Heat-leak detection becomes a black-box test: only the initial and final mean values of passive observables are needed, with no trajectory information and no model of the noise channel.","Because the measurement count grows polynomially with the number of qubits for fixed observable power and shift range, the tests remain feasible where full state or process tomography is not.","Shifted and deformed passive observables detect leaks that the standard Clausius inequality $\\Delta\\langle B\\rangle\\ge0$ and unshifted powers of $B$ miss, so observable choice expands the detectable class.","The tests are self-contained: they need no comparison with a classical simulation of the ideal circuit, which becomes intractable for large devices.","In a constructed example, global passivity detects an environment that leaves resource-theory free-energy constraints unviolated, so the method covers cases where other thermodynamic frameworks fail."],"supporting_citations":[{"why":"Supplies the global passivity constraints and the passive-observable formalism that the heat-leak tests are built on.","marker":"[27]"},{"why":"Introduces passivity deformation, the method used to construct $B_{\\rm def}$ and the second experiment's observable.","marker":"[28]"},{"why":"Provides the shadow-tomography sampling bound used to prove polynomial measurement scaling for the tests.","marker":"[29]"},{"why":"Gives the equivalence between unital maps and mixtures of unitaries plus the majorization condition that underlies the inequality logic.","marker":"[33]"},{"why":"Supplies the exponential tomography lower bound that motivates the polynomial-scaling advantage claimed for the tests.","marker":"[36]"},{"why":"Contains the experimental details, statistical-error analysis, the proof of Eq. (S6), and the constructed heat leaks detectable by any passive observable.","marker":"[35]"},{"why":"Defines the resource-theory free-energy constraints whose failure to detect the environment is contrasted with global passivity.","marker":"[8]"}],"fun_headline_variants":["Thermo violation flags hidden qubit heat leak","Passive observable catches quantum environment leak","Quantum test exposes hidden environment via thermo","Thermodynamic constraint unmasks hidden environment","Hidden environment detected by thermo observable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole detection logic rests on the initial state being exactly the product of thermal states in Eq. (1) after readout correction by the calibrated matrix $M$; if that inversion is biased or the state carries residual correlations, a negative test value could appear without any hidden environment, or a real leak could be missed.","fun_headline_variants_meta":{"raw":{"variants":["Thermo violation flags hidden qubit heat leak","Passive observable catches quantum environment leak","Quantum test exposes hidden environment via thermo","Thermodynamic constraint unmasks hidden environment","Hidden environment detected by thermo observable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000401,"raw_usage":{"total_tokens":2088,"prompt_tokens":937,"completion_tokens":1151,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":1087}},"tokens_in":553,"tokens_out":1151,"duration_ms":11648,"temperature":1.0,"reasoning_tokens":1087,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:26:04.631375+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Prepare a product thermal state on a processor, apply only unital channels (random mixtures of unitaries or depolarizing noise) with the ancilla environment decoupled, and check that every $\\Delta\\langle F_{\\alpha,\\delta}\\rangle$ stays nonnegative; then deliberately perturb the readout-correction matrix $M$ (for example, swap the detection probabilities of one qubit) and see whether a negative test appears. If a calibration error alone can produce a violation, the 'hidden-environment' reading of the inequality is not self-contained.","supporting_citations":[{"cited_title":"Merali, Nature News551, 20 (2017)","cited_arxiv_id":null,"evidence_quote":"Supplies the global passivity constraints and the passive-observable formalism that the heat-leak tests are built on."},{"cited_title":"Experimental detection of microscopic environments using thermodynamic observables","cited_arxiv_id":"1908.08968","evidence_quote":"Provides the shadow-tomography sampling bound used to prove polynomial measurement scaling for the tests."},{"cited_title":"Uzdin and S","cited_arxiv_id":null,"evidence_quote":"Gives the equivalence between unital maps and mixtures of unitaries plus the majorization condition that underlies the inequality logic."},{"cited_title":"Arute, K","cited_arxiv_id":null,"evidence_quote":"Supplies the exponential tomography lower bound that motivates the polynomial-scaling advantage claimed for the tests."},{"cited_title":"Kosloﬀ, Entropy 15(6), 2100 (2013)","cited_arxiv_id":null,"evidence_quote":"Defines the resource-theory free-energy constraints whose failure to detect the environment is contrasted with global passivity."}],"review_version":1}