{"id":"89ec3401-c3dd-4d90-be27-a4a88e787d81","arxiv_id":"2607.08697","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"Measurement incompatibility robustness provides a tight, SDP-computable upper bound on an eavesdropper's guessing probability in prepare-and-measure randomness generation.","lead":"This paper shows that incompatible quantum measurements directly certify randomness in a prepare-and-measure protocol, bounding an eavesdropper's guessing probability via a semidefinite program. A smart generalist might read it because it turns a foundational quantum property—measurement incompatibility—into a practical, computable security certificate for random number generation.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"No significant objection identified. The full-rank assumption is the weakest point but is acknowledged and handled; the SDP formulation and proofs are sound.","rationale":"The reader correctly identified the full-rank assumption as the weakest point. I verified that this is indeed the most load-bearing assumption: the converse direction of Observation 1 fails without it (Eq. 7→8 requires faithfulness), and the quantitative SDP inherits this dependence. However, the assumption is clearly stated, its necessity is acknowledged with citations, and it is standard for semi-device-independent scenarios. The SDP formulation is correct — I checked that the y-independence constraint on Eve's marginal is valid (because e encodes all post-processing as a vector), that the witness constraint correctly probes the same effective measurements in both security and randomness rounds, and that the SDP is a valid relaxation giving an upper bound. The steering translation (Appendix E) is a straightforward variable substitution with matching classical bounds. The quantum-memory extension uses known qubit-simulability bounds appropriately. No internal inconsistency or gap in the logic was found. The reader's ACCEPT verdict with HIGH confidence is appropriate; the correctness risk remains 'unknown' only in the sense that the specific numerical examples have not been independently reproduced, but the theoretical framework is sound.","tokens_in":15894,"tokens_out":6341,"duration_ms":232005,"concrete_test":"Independently solve the SDP (12) for the hollow triangle example (3 qubit MUBs with visibility v∈(1/√3, 1/√2]) using a standard SDP solver (e.g., CVXPY/Mosek). Verify that (i) the guessing probability is strictly less than 1 whenever the witness detects incompatibility (α>0), (ii) the claimed value Pg≈0.888 (with all witness states) is reproduced, and (iii) the bound degrades continuously to Pg→1 as v→1/√3 from above. If the SDP yields Pg=1 for some v with α>0, the quantitative trade-off claim would be compromised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"I traced the central argument carefully: Observation 1's converse direction (Eve guesses perfectly → JM) requires full-rank input states ϱ, which the authors acknowledge and cite as necessary [36,37]. The quantitative SDP in Eq. (12) is a valid relaxation: every physical Eve strategy maps to a feasible point because (i) Eve's post-processed marginal Σ_b G_{e,b|y} is genuinely y-independent when e is a vector encoding guesses for all inputs simultaneously (Eq. 6), (ii) the witness constraint correctly uses Bob's effective measurements B_{b|y}=Σ_e G_{e,b|y}, which are the same in security and randomness rounds, and (iii) the SDP maximizes over a superset of physical strategies, yielding a valid upper bound. The steering equivalence (Appendix E) is a clean change of variables. The quantum-memory extension (Eq. 14) correctly adds qubit-simulability constraints. The full-rank assumption is the genuine soft spot: without it, incompatible measurements can appear compatible on the support of the input states, making the SDP bound trivial (Pg=1) even with genuine incompatibility. But the authors state this explicitly, note that adding white noise to witnesses partially mitigates it, and the assumption is standard for semi-device-independent protocols. I do not find a more serious concern.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This manuscript establishes a quantitative trade-off between measurement incompatibility and randomness generation in a semi-device-independent prepare-and-measure scenario. The central qualitative result (Observation 1) states that, for full-rank input states, the scenario is secure if and only if Bob's effective POVMs are incompatible. The quantitative result (Eq. 12) provides an SDP that bounds Eve's guessing probability given a lower bound on the generalised incompatibility robustness certified by a witness. The authors translate their framework to quantum steering (Appendix E), obtaining a tight connection between steerability and randomness for any finite number of inputs, and extend to an Eve with a qubit quantum memory via dimensional simulability (Eq. 14). The proofs are mathematically careful, with both directions of Observation 1 established explicitly.","tokens_in":16278,"tokens_out":1132,"duration_ms":158663,"significance":"The paper provides a constructive, measurement-theoretic route to randomness certification that is not restricted to two-input scenarios or star-incompatibility structures, improving noise tolerance over prior steering-based protocols. The SDP formulation is parameter-free in the sense that the incompatibility lower bound is obtained from observed data via standard witness duality, not introduced as a free parameter. The steering equivalence (Appendix E) is a clean change of variables that makes the result directly applicable to steering experiments. The quantum-memory extension via qubit simulability is a natural and welcome generalisation. These are genuine advances for semi-device-independent quantum randomness.","major_comments":[{"comment":"Section IV, Observation 1 and Eq. (12): The full-rank assumption on input states is load-bearing for the converse direction of Observation 1 (secure implies incompatible). The authors acknowledge this and cite [36,37], but the quantitative SDP in Eq. (12) inherits this dependence: the witness states must be faithful for the security guarantee to hold. The footnote 3 mentions that adding white noise to the witness allows use of the MUB inequality but notes detection strength may drop. It would strengthen the paper to state more precisely how much the detection strength degrades and whether the SDP bound remains non-trivial after this regularization, at least for the two-MUB example.","section":null},{"comment":"Section V, Example 3 (hollow triangle): The authors report witnessed robustness values and guessing probabilities for dimensions 2, 3, and 4 using a single state per basis, and a guessing probability of 0.888 for the qubit case using all witness states. However, no detail is given on how the witness was constructed for the three-MUB hollow triangle case or which specific witness operators were used. Since this example is the primary illustration of the paper's advantage over prior two-input or star-incompatibility results, a brief specification of the witness (or a reference to where it can be found) would be appropriate.","section":null},{"comment":"Section VI, steering translation: The authors claim a 'tight connection between steerability and randomness generation in a setting using any finite number of measurement inputs.' The mapping in Appendix E is mathematically clean, but the claim of tightness could be stated more precisely. Does tightness mean that the optimal guessing probability in the steering SDP (Eq. E1) equals that of the incompatibility SDP (Eq. 12) for every full-rank state, or does it refer to the qualitative equivalence? Clarifying this would help the reader assess the strength of the steering result.","section":null}],"minor_comments":[{"comment":"Fig. 2 caption: The panels are labeled (a) and (b) but the in-text references in Section V refer to them in order without always specifying which panel. This is minor but could be made explicit.","section":null},{"comment":"Eq. (6): The notation uses a product over tilde-y of p(e_tilde-y | lambda, tilde-y), which is correct for deterministic post-processing but could be confusing on first read. A brief clarifying sentence that this encodes a deterministic strategy vector would help.","section":null},{"comment":"Section V, Example 4: The guessing probability of 0.924 for the qubit-memory Eve is stated without specifying the corresponding incompatibility robustness value or whether this is at maximal robustness. Adding the context would make the number more interpretable.","section":null},{"comment":"Reference [48] in the note added: The distinction drawn between the authors' requirement (Eve guesses for some full-rank state unknown to her) and the related work's requirement (guessing for a collection of input states) is somewhat subtle. A slightly more explicit comparison in the main text or a footnote would help readers appreciate the difference.","section":null},{"comment":"Typographical: 'PREP ARE-AND-MEASURE' and 'INCOMP A TIBILITY' in section headers appear to have spacing artifacts from the source.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is well-structured and the core results are sound. The full-rank assumption is the genuine soft spot but is handled honestly and is standard for semi-device-independent protocols. The main improvement needed is clarification of the witness construction in Example 3 and tighter discussion of the regularization cost mentioned in footnote 3. I do not see a deeper problem."},"author_rebuttal":null,"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper gives a quantitative trade-off between measurement incompatibility (specifically the generalised robustness) and Eve's guessing probability in a semi-device-independent prepare-and-measure scenario. The qualitative connection was known from prior work [29-31], but the quantitative SDP formulation (Eq. 12), the extension to arbitrary measurement structures beyond two-input or star-incompatibility, and the clean steering translation are genuinely new. This deserves a serious referee. The core argument holds up. Observation 1 (secure iff incompatible, for full-rank states) is proved in both directions with explicit construction. The Naimark dilation argument for the JM-to-insecure direction is clean. The converse uses the full-rank assumption on ϱ to promote a trace inequality to operator equality (Eq. 7 to Eq. 8), which is correct and necessary. The central SDP (Eq. 12) correctly encodes pairwise joint measurability constraints on Eve's post-processed POVMs, and the objective maximizes over a superset of physical strategies, so it gives a valid upper bound on Pg. The steering equivalence in Appendix E is a straightforward change of variables and checks out. The hollow triangle example (Example 3) is a nice illustration that randomness can be certified even when all measurement pairs are jointly measurable, which is a real advantage over prior steering-based protocols. The quantum memory extension (Eq. 14) is more preliminary — it replaces the JM constraint with qubit-simulability and relies on known qubit bounds for MUBs to get a single SDP, but the general case would require more computational machinery. The soft spot is the full-rank assumption, and the authors know it. Without it, incompatible measurements can look compatible on the support of the input states, making the SDP bound trivial. They cite the relevant counterexamples [36, 37] and note that adding white noise to witnesses partially mitigates this at the cost of detection strength. This is honest and the assumption is standard for semi-DI protocols, but it does mean the abstract's claim of a protocol for 'any set of incompatible measurements' is slightly oversold — it's any set probed by a faithful witness. The examples are illustrative rather than exhaustive. The numerical values for the hollow triangle (robustnesses 0.082-0.125, guessing probabilities 0.888-0.995) are plausible but I could not independently verify them. This is a theory paper for people working on semi-device-independent quantum information, measurement theory, or steering. It is not a breakthrough but it is a careful, well-constructed contribution that fills a real gap between the qualitative incompatibility-randomness connection and a usable quantitative protocol. Recommend accept for peer review.","headline":"Solid theoretical contribution linking incompatibility robustness to randomness generation; the full-rank assumption is the real but acknowledged limitation.","tokens_in":16619,"tokens_out":1586,"would_cite":true,"duration_ms":39613,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Dd","03.67.Hk","03.65.Ta"],"model":"glm-5.2","headline":"Incompatible measurements bound eavesdroppers and certify randomness","keywords":["measurement incompatibility","randomness generation","prepare-and-measure","semidefinite programming","quantum steering","joint measurability","incompatibility robustness","device-independent quantum information"],"falsifier":"A concrete counterexample would be a set of measurements that is genuinely incompatible (positive incompatibility robustness) and a set of full-rank input states for which the SDP in Eq. (12) still yields Eve's guessing probability equal to 1, meaning the certified incompatibility does not translate into certified randomness.","tokens_in":16201,"feed_emoji":"","tokens_out":1107,"duration_ms":122898,"temperature":0.7,"pith_summary":"The paper establishes that measurement incompatibility—the impossibility of jointly measuring two or more quantum observables—is not merely a foundational curiosity but a directly quantifiable resource for generating secret randomness. In a prepare-and-measure scenario where Alice sends quantum states to an untrusted receiver Bob while an eavesdropper Eve intercepts messages, the authors prove that the scenario is secure (Eve cannot perfectly predict Bob's outcomes) if and only if Bob's effective measurements are incompatible, provided the input states have full rank. They then make this qualitative statement quantitative: the generalised robustness of incompatibility, a geometric measure of how much noise must be added to a set of measurements before they become jointly measurable, upper-bounds Eve's guessing probability through a semidefinite program. Any incompatibility witness—a set of trusted test states whose statistics reveal incompatibility—doubles as a randomness certificate, yielding an explicit protocol that works for any finite number of measurement inputs and any incompatible measurement set.","feed_headline":"Incompatible quantum measurements bound eavesdroppers and certify randomness","feed_subtitle":"A semidefinite program converts any measurement incompatibility into a quantitative limit on an eavesdropper's guessing power, yielding a通用型","key_machinery":"The load-bearing object is the semidefinite program (SDP) in Eq. (12), which maximises Eve's guessing probability subject to two constraints: (1) Bob's effective measurements must witness at least a threshold amount of incompatibility robustness, and (2) Eve's joint POVMs must be pairwise jointly measurable with Bob's effective POVMs. The incompatibility witness—trusted positive semidefinite operators whose expectation values are bounded for all jointly measurable measurement sets—provides the lower bound on incompatibility that feeds into the SDP, while the pairwise joint measurability constraint encodes the structural limit on Eve's strategies.","core_discovery":"The central mechanism is a structural identity between Eve's information-gathering capability and the mathematical structure of joint measurability. Eve's attack is fully characterised by POVMs that are pairwise jointly measurable with Bob's effective measurements: she can perfectly guess Bob's outcome for a given input if and only if her marginal measurement reproduces his. This means incompatibility of Bob's effective measurements is the exact condition preventing perfect prediction. The authors convert this qualitative equivalence into a quantitative bound by using the generalised incompatibility robustness, which measures the distance from a measurement set to the jointly measurable set.","pith_inferences":["If the full-rank assumption on input states could be relaxed—for instance, by using dimensional witnesses or self-testing techniques—the protocol would become more device-independent and potentially more experimentally practical, since preparing and verifying full-rank states adds overhead.","The connection between incompatibility robustness and state discrimination tasks suggests that optimal randomness-generation protocols could be designed by choosing measurement sets that are maximally hard to discriminate, creating a design principle for randomness sources.","Extension to continuous-variable systems seems feasible given that incompatibility robustness has infinite-dimensional counterparts, potentially enabling randomness certification from quadrature measurements or homodyne detection."],"forward_implications":["Any set of incompatible measurements, regardless of structure or number of inputs, can be converted into a randomness generation protocol by selecting appropriate test states and solving a single SDP.","The framework extends to quantum steering: steerability of a state assemblage and randomness certification are tightly linked for any finite number of measurement inputs, improving noise tolerance over prior two-input or star-incompatibility restrictions.","An eavesdropper with a quantum memory can be bounded by replacing joint measurability with dimensional simulability, allowing the framework to handle more powerful adversaries.","Incompatibility monogamy relations emerge naturally: higher incompatibility robustness means fewer compatible strategies for Eve, suggesting a resource-theoretic trade-off between incompatibility and eavesdropper capability."],"fun_headline_variants":["Measurement incompatibility quantifies how much randomness eavesdroppers can't predict","Robustness of incompatible measurements bounds Eve's guessing power in prepare-and-measure","Any incompatible measurements generate randomness with certified eavesdropper limits","Incompatibility robustness certifies randomness by bounding Eve's joint measurability","Steerability and randomness generation linked through measurement incompatibility bounds"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The security proof requires that the input states sent by Alice have full rank, meaning they are supported on the entire Hilbert space. Without this, there exist incompatible measurements that look jointly measurable on the subspace the states actually probe, and Eve could exploit this gap to perfectly guess outcomes despite apparent incompatibility.","fun_headline_variants_meta":{"raw":{"variants":["Measurement incompatibility quantifies how much randomness eavesdroppers can't predict","Robustness of incompatible measurements bounds Eve's guessing power in prepare-and-measure","Any incompatible measurements generate randomness with certified eavesdropper limits","Incompatibility robustness certifies randomness by bounding Eve's joint measurability","Steerability and randomness generation linked through measurement incompatibility bounds"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":582,"prompt_tokens":484,"completion_tokens":98,"prompt_tokens_details":null},"tokens_in":484,"tokens_out":98,"duration_ms":49321,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T02:46:54.564992+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"A concrete counterexample would be a set of measurements that is genuinely incompatible (positive incompatibility robustness) and a set of full-rank input states for which the SDP in Eq. (12) still yields Eve's guessing probability equal to 1, meaning the certified incompatibility does not translate into certified randomness.","supporting_citations":[],"review_version":2}