{"id":"cff2a184-7d7c-4040-a2b4-281865703aed","arxiv_id":"2608.05240","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Quantum random-access codes let a one-qubit-per-weight quantizer retrieve context-dependent signs, strictly improving reconstruction risk over shared-sign one-bit PTQ when context-wise optimal signs disagree.","lead":"A proposal to store a neural network's quantized weights in quantum states so different deployment contexts can read out different binary sign patterns, instead of sharing one sign pattern. The paper proves this beats a specific classical one-bit baseline when contexts want different signs, though only under an idealized memory model.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem C.7's classical baseline excludes affine zero-point decoders; against that class the separation reverses for M=2, so the 'one qubit beats one bit' claim is not supported.","rationale":"The reader's weakest assumption (fresh-copy readout) is an explicitly acknowledged hardware-model limitation that does not affect the ideal S=∞ separation; the affine-baseline gap refutes the generality of the ideal separation itself by a concrete zero-risk classical construction, so it is more load-bearing for the central claim. The formal theorem as stated is correct within its chosen baseline, so this is a scope/claim problem rather than an internal inconsistency. It reinforces the reader's CONDITIONAL verdict for an independent reason: the paper needs either a comparison against affine one-bit PTQ or a re-scoped title and abstract. The reader's rationale already noted the affine omission, but chose fresh-copy as the weakest assumption, hence partial agreement.","tokens_in":40961,"tokens_out":13602,"duration_ms":148837,"concrete_test":"On the two-context symmetric family Σ±=I±r(J-I) with M=2 and any nonzero row w, evaluate the affine one-bit risk min_{c∈{±1}2, ατ, βτ} Στ πτ (w−ατ c−βτ 1)Στ(w−ατ c−βτ 1)^T. For c=(1,−1) the 2×2 system ατ c+βτ 1=w is solvable, so the optimum is 0, while ideal QRAQ risk from Corollary C.9 is (1−r)(w1^2+w2^2)/2>0 for r∈(0,1). This settles: the claimed separation holds only against the no-bias signed-scale subclass, not against the full one-bit class the paper itself acknowledges.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The main separation Theorem C.7 (Eq. 56) optimizes a shared sign c and per-context scales ατ with no per-context bias. Yet Theorem C.5 explicitly identifies affine one-bit quantizers with a context-dependent zero-point as the classical simulation of non-symmetric fixed-readout decoders, and the paper never compares QRAQ to that class. For M=2 the omission is decisive: choose shared c=(1,-1); then span{c, 1}=R^2, so per context τ solve ατ c + βτ 1 = w_i,:, giving zero classical reconstruction risk for every row. In the same setting ideal QRAQ risk is positive for r>0 (Corollary C.9, Eq. 73), so QRAQ is strictly worse than an affine one-bit baseline. Thus the central separation is an artifact of restricting the classical side to no-bias signed per-row scales; no theorem in the paper bounds the affine class, and the title/abstract's unqualified 'one qubit beats one bit' overreaches. The paper should either prove advantage over affine one-bit PTQ (likely false) or restate all claims relative to the no-bias subclass.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Quantum Random Access Quantization (QRAQ), a scheme in which each weight entry stores one logical QRAC register and the deployment context selects a Pauli measurement, producing an unbiased context-dependent sign estimate with an explicit shot-noise penalty. The main formal result (Theorem C.7) proves that, for signed per-row scales and no zero-point, ideal QRAQ per-row risk is never worse than the best shared-sign classical one-bit risk, and is strictly better exactly when the context-wise optimal sign sets have no common representative modulo global sign. Finite-shot and depolarizing-noise thresholds are derived, and the sign-disagreement condition is shown necessary for any finite-shot advantage. The paper also provides a scale-granularity boundary, a finite-sample certificate, extensions to multi-qubit QRACs, correlated noise, and general Pauli noise, together with simulator experiments that match the closed-form predictions.","tokens_in":41173,"tokens_out":9664,"duration_ms":101040,"significance":"Within the deliberately narrow class of no-bias signed per-row shared-sign one-bit quantizers, the formal results are sound and clearly presented. The proofs are self-contained, the shot-noise and noise thresholds are explicit, and the finite-sample certificate is a practically useful addition. The use of measurement incompatibility as the operative resource is articulated well, and the experimental section is honest simulator validation rather than an overclaimed hardware benchmark. However, the headline claim that one qubit beats one bit is not supported against affine one-bit quantizers with a zero-point, a class that the paper's own Theorem C.5 identifies as the classical simulation of non-symmetric fixed-readout decoders; in the two-dimensional case that affine class attains zero classical risk while ideal QRAQ has positive risk. The fresh-copy readout model also makes the one-qubit-per-weight memory claim conditional in a way the title does not convey. The contribution is therefore a valid separation result for a well-defined subclass, but it is substantially narrower than advertised.","major_comments":[{"comment":"The central separation is stated only against shared-sign signed per-row scales with no zero-point, yet Theorem C.5 explicitly introduces affine one-bit quantizers with a context-dependent zero-point as the classical simulation of non-symmetric fixed-readout decoders. The affine class is therefore within the paper's own logical landscape and is a natural one-bit PTQ baseline. For M=2 and K=2, choose the shared sign c=(1,-1); for any row w=(w1,w2) and any context tau, the scalars alpha_tau=(w1-w2)/2 and beta_tau=(w1+w2)/2 satisfy alpha_tau c + beta_tau 1 = w, so the affine shared-sign baseline attains zero reconstruction risk in every context. In the same two-dimensional symmetric-covariance setting, Corollary C.9 gives positive ideal QRAQ risk whenever r>r0, for example w=(1,3) and r=0.8 gives ideal QRAQ risk 1. Thus the unqualified claim in the title and abstract that one qubit beats one bit is false for the affine class, and no theorem in the paper bounds that class. The claims should be restated relative to the no-bias signed per-row subclass, or a separate treatment of the affine baseline should be added.","section":"Theorem C.7; Eq. (56); Corollary C.9, Eq. (73)"},{"comment":"The fresh-copy readout model is load-bearing for the 'one qubit per weight' resource claim. If state re-preparation is unavailable, serving S shots or K contexts requires S or K physical qubit preparations per weight, which erases the one-qubit memory advantage. If re-preparation is available, the device must have access to the calibrated context-specific sign matrices B^(tau) to prepare each QRAC state; that is exactly the information QRAQ is claimed to compress. The paper acknowledges this in Assumption 4.1 and Section 7, but the title, abstract, and main-takeaway box do not condition the advantage claim on this model. The abstract and conclusion should state that the result is a per-query logical readout separation under the fresh-copy/re-preparation model, not a physical memory-size advantage.","section":"Assumption 4.1; Section 7"}],"minor_comments":[{"comment":"The phrase 'or equivalently on a device that can re-prepare the same QRAC state before each shot' is statistically equivalent but not operationally or resource equivalent; consider replacing 'equivalently' with 'alternatively' and explicitly discussing the resource difference.","section":"Assumption 4.1"},{"comment":"The limitations list should include the affine zero-point per-row class as a class over which QRAQ has no certified advantage; the current list mentions per-entry scales but omits the affine one-bit class highlighted by Theorem C.5.","section":"Section 7, Limitations"},{"comment":"The observation that the sufficient threshold of Corollary C.8 can never certify advantage at S=1 in the symmetric two-context family is interesting and would be more visible if mentioned near Eq. (16) in the main text rather than only in the appendix.","section":"Corollary C.10"}],"recommendation":"major_revision","confidential_remarks":"The authors are transparent about many limitations, which I appreciate. The main concern is that the headline claim is broader than the theorems actually prove. If the authors narrow the title and abstract and add a short appendix or section treating the affine one-bit baseline, the formal core is publishable. I recommend major revision rather than rejection because the formal statements are internally consistent and the restricted separation result is correct."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is a carefully written theory paper, but the headline claim is not supported as stated. The formal separation is genuine and the proofs look correct, but only against classical one-bit quantizers with signed per-row scales and no per-context zero-point. The paper itself introduces affine one-bit quantizers (sign plus bias) as the classical simulation class for non-symmetric fixed readouts, yet never compares QRAQ to them. That omission is decisive for M=2: choose a shared sign c=(1,-1); per context, span{c, 1} is the whole row space, so an affine baseline reproduces every row exactly and has zero reconstruction risk. Ideal QRAQ has positive risk in the same setting (Corollary C.9). So against the affine class the advantage reverses. This is not a trivial edge case; it means the title and abstract promise more than the theorems deliver.\n\nWhat the paper does well: Theorem C.7 is a clean structural result — when context-wise optimal signs have no common representative, QRAQ strictly improves over the shared-sign no-bias baseline, with explicit finite-shot and noise thresholds. The fixed-readout no-go (Theorem C.5) is a useful sanity check that isolates measurement incompatibility as the resource. The closed-form two-dimensional gap, scale-granularity boundary, and finite-sample certificate are concrete and appear correct. The simulator experiments are honest and match the formulas. It is not sloppy work.\n\nThe second soft spot is Assumption 4.1 (fresh-copy readout). The paper discloses it clearly, but it does real work: without re-preparation, serving S shots means S physical qubit copies per weight, which erases the one-qubit-per-weight memory claim. This is a logical-model contribution, not a deployment claim. Fine for a theory paper, but it further narrows the title's promise.\n\nWho is this for? People working on the theory of low-bit quantization or on QRAC applications in ML. It is not going to change PTQ practice soon. The proofs deserve referee time, but the authors need to re-scope: prove or disprove advantage over affine one-bit (likely negative), or state the separation only for the no-bias subclass and adjust the title and abstract accordingly.\n\nMy recommendation: send it to peer review, conditional on a serious revision that either restricts the claims to the class where the theorem holds or adds the missing affine-class analysis.","headline":"The separation proof is real but only against no-bias signed per-row one-bit baselines, not the affine one-bit class the paper itself identifies, so the 'one qubit beats one bit' headline overreaches.","tokens_in":41714,"tokens_out":3956,"would_cite":true,"duration_ms":40760,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A qubit can beat a classical bit in one-bit model quantization when contexts demand different signs.","keywords":["post-training quantization","one-bit quantization","quantum random access code","measurement incompatibility","context-dependent weight signs","shared-sign constraint","reconstruction risk","shot noise"],"falsifier":"On the two-context symmetric covariance family with $w=(1,0)$, $S=1$, and $\\eta=1$, compute the row risk of a physical single-qubit QRAQ implementation and compare to the one-bit shared-sign baseline; the paper predicts a strictly lower QRAQ risk exactly for anisotropy $r$ in $((\\sqrt{5}-1)/2, 1)$. If the measured risk gap is non-positive in that interval, or if the gap fails to vanish below that threshold, the central separation is refuted. Alternatively, run the same comparison on real attention-head activation covariances and verify that rows with positive ideal gap correspond to empty intersection of per-context sign optima.","tokens_in":40718,"feed_emoji":"⚛️","tokens_out":9394,"duration_ms":81623,"temperature":0.7,"pith_summary":"This paper asks whether one logical qubit per weight can outperform one classical bit in the most extreme form of model compression: storing each weight only as a sign. It introduces Quantum Random Access Quantization (QRAQ), which encodes several context-specific sign matrices into one quantum state and retrieves the sign for the current deployment context by choosing a matching Pauli measurement. For signed per-row scales, the paper proves that ideal QRAQ attains the average of the context-wise binary optima, while any classical one-bit quantizer must commit to a single sign vector shared across all contexts. The row-wise reconstruction-risk gap is strictly positive exactly when the context-wise optimal sign sets have no common representative, and it survives finite shots and depolarizing noise when the ideal margin exceeds an explicit variance inflation. The result identifies measurement incompatibility, rather than quantization error alone, as a resource that could be used if future hardware supports calibrated re-preparation of qubit states.","feed_headline":"One qubit beats one bit when contexts want different signs","feed_subtitle":"A qubit stores context-specific weight signs; the gain appears exactly when contexts disagree on signs.","key_machinery":"The central object is a quantum random-access code (QRAC) used as a memory cell per weight entry. For two contexts, the cell is a single qubit in state $\\rho_{ij} = (I + B^{(1)}_{ij} X/\\sqrt{2} + B^{(2)}_{ij} Z/\\sqrt{2})/2$, whose $X$ and $Z$ measurements return the two possible signs in expectation; for $K$ contexts, the construction uses $n = \\lceil(K-1)/2\\rceil$ qubits and a Jordan–Wigner family of pairwise anticommuting Pauli observables, with Bloch radius $1/\\sqrt{K}$ tight for positivity. The mechanism: a context label selects the matched Pauli observable, and averaging $S$ fresh-copy readouts yields an unbiased estimate of the context-specific sign with an explicit shot-noise term $\\nu_K(\\eta)/S$, where $\\nu_K(\\eta) = K/\\eta^2 - 1$. The proof machinery is the bias–variance decomposition of the layer-wise quadratic reconstruction risk, reduced row-wise to a comparison of 'min after summing' versus 'sum of mins.'","core_discovery":"At the heart of the paper is a row-wise comparison. For a weight row $w$ and context covariance $\\Sigma_\\tau$, the best one-bit fit to a sign vector $b$ is $J^*_\\tau(b) = w \\Sigma_\\tau w^T - (b \\Sigma_\\tau w^T)^2/(b \\Sigma_\\tau b^T)$. The classical shared-sign row risk is $\\min_{c \\in \\{\\pm1\\}^M} \\sum_\\tau \\pi_\\tau J^*_\\tau(c)$, while ideal QRAQ attains $\\sum_\\tau \\pi_\\tau \\min_b J^*_\\tau(b)$. The paper proves (Theorem C.7) that the difference $\\Delta_\\infty = \\min_{c} \\sum_\\tau \\pi_\\tau J^*_\\tau(c) - \\sum_\\tau \\pi_\\tau \\min_b J^*_\\tau(b)$ is nonnegative, and strictly positive exactly when the per-context optimal sign sets $\\mathcal{M}_\\tau = \\arg\\min_b J^*_\\tau(b)$ have empty intersection modulo the global sign symmetry $b \\sim -b$. For finite shots, QRAQ's unbiased estimate of each context sign carries variance $\\mu^2/S \\cdot (K/\\eta^2 - 1)$ under depolarizing noise, so the strict row-wise separation survives when $S$ exceeds the explicit inflation ratio in Eq. (16). The result turns the shared-sign constraint of one-bit PTQ into a quantifiable gap that can be certified from calibration data.","pith_inferences":["The advertised one-qubit-per-weight memory claim is conditional on the fresh-copy readout model: without re-preparation or parallel instantiation, serving $K$ contexts or $S$ shots consumes $K$ or $S$ physical qubit copies per weight, which the paper itself lists as a limitation.","The method requires the encoder to know the context-wise sign optima at calibration time, so it compresses information that a classical device could also store as $K$ sign bits; the quantum advantage is a reconstruction-risk statement under a fixed memory cell count, not a communication advantage.","A practical deployment path suggested by the certificate is to evaluate the sign-disagreement margin on calibration data first and spend shot budget only on rows where the ideal gap clears the noise threshold, making the advantage testable before any quantum hardware exists.","If confirmed on real attention-head or mixture-of-experts activation statistics, the result would reframe one-bit PTQ as a measurement-selection problem over a fixed quantum memory, connecting quantization to quantum memory design."],"forward_implications":["Under signed per-row scales, the ideal QRAQ row risk never exceeds the shared-sign classical optimum, and it is strictly lower precisely when the per-context optimal sign sets for that row have empty intersection modulo global sign symmetry.","The finite-shot separation is governed by a computable threshold: a row keeps its advantage when the shot count $S$ exceeds the variance inflation $\\nu_K(\\eta) \\sum_\\tau \\pi_\\tau T_\\tau (\\mu^*_\\tau)^2 / \\Delta_\\infty^i$, with $\\nu_K(\\eta) = K/\\eta^2 - 1$.","For scale classes coarser than or equal to signed per-row scaling, the separation transfers automatically; for per-column, row-times-column, or group scaling, no universal gap exists and the margin must be certified directly on the instance.","In the resource-fair comparison (one qubit versus one classical bit, $S = 1$), quantum advantage persists in the two-context symmetric covariance family when the anisotropy $r$ exceeds $(\\sqrt{5}-1)/2$ for $w = (1,0)$.","Empirical gap certificates: with bounded activations and nondegenerate denominators, sample sizes satisfying $M \\varepsilon_\\Sigma \\le \\lambda_0/2$ certify a positive population gap with probability at least $1-\\delta$ via the stated confidence radius."],"supporting_citations":[{"why":"Introduces quantum random access codes, the encoding primitive QRAQ uses to store multiple context signs in one state.","marker":"[Ambainis et al., 2002]"},{"why":"Provides the information-theoretic bound that a single qubit cannot reveal more than one requested bit, motivating context-matched measurements.","marker":"[Nayak, 1999]"},{"why":"Establishes the link between QRAC success and measurement incompatibility, the resource the paper claims drives the advantage.","marker":"[Carmeli et al., 2020]"},{"why":"Supplies the layer-wise quadratic reconstruction objective that both classical and QRAQ quantizers optimize.","marker":"[Frantar et al., 2023]"},{"why":"Activations-aware weight quantization that motivates per-context activation covariances.","marker":"[Lin et al., 2024]"},{"why":"BitNet, the shared-sign one-bit LLM baseline whose per-row signed scales define the classical feasible set.","marker":"[Wang et al., 2023]"},{"why":"The 1-bit LLM regime that QRAQ targets, establishing the practical one-bit-per-weight setting.","marker":"[Ma et al., 2024]"},{"why":"Gives the Jordan–Wigner construction of pairwise anticommuting observables used to encode $K$ contexts on $\\lceil(K-1)/2\\rceil$ qubits.","marker":"[Jordan and Wigner, 1928]"}],"fun_headline_variants":["Quantum signs beat binary when contexts disagree","QRAQ: context-matched qubits outdo shared one-bit weights","Measurement incompatibility drives PTQ quantum advantage","One qubit beats one bit when sign needs vary","Quantum encoding of signs wins over fixed binary PTQ"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All claims of one-qubit memory advantage rely on the fresh-copy readout model (Assumption 4.1): each measurement shot needs a newly prepared copy of the same QRAC state; if a physical device cannot re-prepare or parallel-instantiate states, the one-qubit-per-weight memory saving evaporates.","fun_headline_variants_meta":{"raw":{"variants":["Quantum signs beat binary when contexts disagree","QRAQ: context-matched qubits outdo shared one-bit weights","Measurement incompatibility drives PTQ quantum advantage","One qubit beats one bit when sign needs vary","Quantum encoding of signs wins over fixed binary PTQ"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1343,"prompt_tokens":1059,"completion_tokens":284,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":675,"completion_tokens_details":{"reasoning_tokens":209}},"tokens_in":675,"tokens_out":284,"duration_ms":4035,"temperature":1.0,"reasoning_tokens":209,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T17:26:30.915603+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On the two-context symmetric covariance family with $w=(1,0)$, $S=1$, and $\\eta=1$, compute the row risk of a physical single-qubit QRAQ implementation and compare to the one-bit shared-sign baseline; the paper predicts a strictly lower QRAQ risk exactly for anisotropy $r$ in $((\\sqrt{5}-1)/2, 1)$. If the measured risk gap is non-positive in that interval, or if the gap fails to vanish below that threshold, the central separation is refuted. Alternatively, run the same comparison on real attention-head activation covariances and verify that rows with positive ideal gap correspond to empty intersection of per-context sign optima.","supporting_citations":[{"cited_title":"Optimal lower bounds for quantum automata and random access codes","cited_arxiv_id":null,"evidence_quote":"Provides the information-theoretic bound that a single qubit cannot reveal more than one requested bit, motivating context-matched measurements."}],"review_version":1}