{"id":"35354296-f383-4556-92af-72431099b5ed","arxiv_id":"2608.09625","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A Jacobian Gram matrix framework quantifies how seven density-matrix parameterizations affect diffusion-based quantum state tomography, showing isometry and constraint satisfaction are competing and that conditioning alone does not predict end-to-end performance.","lead":"This paper studies how different mathematical ways of representing a quantum state, called parameterizations, affect diffusion-based quantum state tomography, and introduces two metrics based on the Jacobian Gram matrix to measure coordinate conditioning.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's validation claim is contradicted by the paper's own matched 3-qubit comparison: perfectly conditioned Bloch (1.0/1.0) is the worst performer, while worse-conditioned Hermitian direct wins.","rationale":"The reader's conditional verdict is well aligned with my read: the paper provides a useful, likely reproducible geometric calibration, but the end-to-end validation does not support the abstract's strong causal claim. My most load-bearing concern differs slightly from the reader's stated weakest assumption. The reader emphasizes that the Jacobian Gram is calibrated at a fixed set of interior states, which may not represent the noise-perturbed geometry along a diffusion trajectory. That is a real concern about the metric's validity for diffusion training. However, the more decisive problem is internal to the paper: even granting the calibration numbers at face value, the paper's own matched 3-qubit training comparison (Table 8) shows the best-conditioned parameterization (Bloch, 1.0×/1.0×) performing far worse than a worse-conditioned one (Hermitian direct, 2.0×/2.0×) under a uniform protocol with no CFG. This is not a subtle confound; it is a direct counterexample to the abstract's sentence that better-conditioned parameterizations converge faster and achieve higher fidelity. The 2-qubit comparison is likewise confounded by unequal learning rates, and the paper's cross-validation in §4.5 only varies λ_meas, not lr. The paper's own 'paradox' discussion and open question in §5.5 are honest signals that the mechanism is not understood. I do not recommend rejection because the geometric calibration, the scaling analysis, and the CFG boundary-effect diagnostics are independently valuable, and the overclaim is cleanly removable. The verdict should remain conditional, with the condition being a revision of the central claim and a fully matched validation, which is exactly what the reader requested. My agreement is 'partial' because the reader's formal weakest assumption (fixed-state calibration) is not the same as my strongest concern (internal contradiction in the matched comparison), even though both point to the need for more careful validation.","tokens_in":11456,"tokens_out":5919,"duration_ms":56496,"concrete_test":"Complete the 3-qubit end-to-end evaluation promised in Section 4.6: report test-set fidelity versus shot count for Bloch, Hermitian direct, and Cholesky under the identical matched protocol (same lr = 2×10^-4, batch = 128, λ_meas = 0.0, seeds) with per-state error bars and K samples per state. If Bloch, the perfectly conditioned parameterization, remains worse than Hermitian direct in this fully matched no-CFG evaluation, the abstract's claim that better-conditioned parameterizations converge faster and achieve higher fidelity is false for the cleanest comparison in the paper. The paper should then be revised to claim only that conditioning can matter among parameterizations with comparable domain geometry, and that bounded-domain effects dominate the Bloch-versus-unbounded comparison.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 1.3 assert that end-to-end training confirms better-conditioned parameterizations converge faster and reach higher fidelity in the absence of CFG. This is contradicted by the paper's own near-matched 3-qubit validation. Table 3 reports Bloch as perfectly conditioned (κ_spec = κ_diag = 1.0×) and Hermitian direct as 2.0×/2.0×; Table 8, using the same lr = 2×10^-4, batch = 128, and λ_meas = 0.0 for all models, shows Bloch at 0.4507 versus Hermitian at 0.7987 validation fidelity after 300 epochs. Thus the best-conditioned parameterization is the worst performer by a large margin under a matched protocol and without CFG. At 2 qubits the same reversal appears (Table 7, λ_meas = 0.0: Bloch Fproj = 0.9188 vs Herm B = 0.9264) and is additionally confounded by different learning rates (Appendix C: Bloch 2×10^-3, Herm B 4×10^-3), so the 2-qubit comparison cannot rescue the claim. The paper honestly labels this an 'isometry–performance paradox' (§4.2) and leaves the mechanism open (§5.5), but that means the central causal statement in the abstract is not supported. For the claim to hold, SDR/DA would have to predict ranking in the decisive matched comparison; they do not. The calibration atlas may remain useful as descriptive geometry, but the headline validation conclusion needs revision or removal.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper calibrates the local Euclidean geometry of seven density-matrix parameterizations for diffusion-based quantum state tomography by computing the Jacobian Gram matrix J^T J and reporting two conditioning metrics, spectral dynamic range and diagonal anisotropy, at 2- and 3-qubit scales. It then trains diffusion models on three of these parameterizations and claims that end-to-end training confirms that better-conditioned parameterizations converge faster and reach higher fidelity without classifier-free guidance. The calibration procedure is clearly specified and the paper is transparent about protocol differences, but the central validation claim is contradicted by the paper's own matched comparisons, leaving the contribution as a descriptive geometric atlas rather than a validated selection rule.","tokens_in":11763,"tokens_out":6701,"duration_ms":56975,"significance":"If the calibration atlas is robust, it provides a practical diagnostic for comparing density-matrix parameterizations and quantifies known scaling problems such as the exponential map's severe conditioning degradation. The paper's strengths include an explicit finite-difference calibration protocol, reproducible state sampling, honest reporting of the 'isometry-performance paradox' and protocol confounds, and a falsifiable per-coordinate sigma-data normalization proposal. However, because the headline causal claim is not supported by the matched experimental data, the atlas is currently a descriptive tool plus an open problem; its value as an engineering guideline depends on resolving the contradictions identified below.","major_comments":[{"comment":"The central claim that 'better-conditioned parameterizations converge faster and achieve higher fidelity in the absence of CFG' is contradicted by the matched 3-qubit comparison. Table 3 reports Bloch/Gell-Mann as perfectly conditioned (κ_spec = κ_diag = 1.0×) and Hermitian direct as 2.0×/2.0×; Table 8, with identical lr = 2×10^-4, batch = 128, λ_meas = 0.0 and no CFG, shows Bloch at 0.4507 and Hermitian at 0.7987 validation fidelity after 300 epochs. The best-conditioned parameterization is the worst performer by a large margin. The same reversal appears at 2 qubits in Table 7 under λ_meas = 0.0 (Bloch 0.9188 vs. Herm B 0.9264), though that comparison is confounded by learning rate. The paper labels this an 'isometry-performance paradox' (§4.2) and leaves the mechanism open (§5.5), which is honest, but it means the abstract's validation sentence is unsupported. The sentence should be revised or removed, and the conclusion should state that conditioning alone does not predict no-CFG performance.","section":"Abstract; §1.3; §4.2; Tables 3 and 8"},{"comment":"Section 4.4 and Figure 6 report a Pearson correlation ρ = -0.87 between κ_diag and final validation fidelity with n = 3. Three points cannot establish a correlation, and the two points that are directly compared under a matched protocol (Bloch and Herm B) have the opposite sign; the negative value is driven by the Cholesky point's large κ_diag. This 'geometry-performance correlation' is therefore not evidence for the claimed monotonic relationship and should be removed or replaced with a scatter of many independent runs.","section":"§4.4; Figure 6"},{"comment":"The 2-qubit no-CFG ordering remains confounded by learning rate: Appendix C gives lr = 2×10^-3 for Bloch and 4×10^-3 for Herm B, and §5.5 acknowledges this. The two-by-two cross-validation in §4.5 varies λ_meas but keeps the same learning-rate difference, so it cannot establish that the parameterization, rather than optimization speed, drives the Herm B advantage. A matched learning-rate run, or a learning-rate sweep showing that the ordering is stable, is required before any performance claim can be attributed to the parameterization.","section":"§4.5; Appendix C; §5.5"},{"comment":"The calibration protocol evaluates J^T J at 30 fixed interior states, and §5.5 notes that dynamic calibration along the diffusion trajectory could reveal time-dependent effects. Since §5.4 reports out-of-domain fractions of 0.55–0.99 for sampled outputs, the effective geometry encountered during training and sampling may differ substantially from the interior atlas. The selection guidelines in §3.5–§3.7 rely on the static values; until the Gram spectrum is evaluated on noise-perturbed coordinates, or a trajectory-dependent calibration is provided, the atlas supports a descriptive comparison of charts at interior states but not a validated engineering rule for diffusion training.","section":"§2.4; §5.4; §5.5"},{"comment":"The per-coordinate sigma-data prescription σ_data,ii ∝ sqrt(G_ii) and Algorithm 7 are presented as actionable implications of the framework, but no experiment compares this schedule against the standard global EDM sigma-data calibration. Without an ablation showing that the coordinate-aware schedule improves convergence or final fidelity, this recommendation is not validated by the paper's experiments.","section":"§5.2; Algorithm 7; Eq. (4)"}],"minor_comments":[{"comment":"The text refers to 'Table 1 summarizes the seven parameterizations' but the calibration results are in Table 2; the cross-reference should be corrected.","section":"§3.2"},{"comment":"The caption of Figure 5 says Herm B converges faster 'despite worse nominal isometry', which is inconsistent with the abstract's causal claim; if the abstract is revised, the caption and related text in §4.2 should be revised consistently.","section":"Figure 5"},{"comment":"The reported σ_data values do not specify whether they are scalars or per-coordinate vectors; this should be clarified to match the per-coordinate prescription in §5.2.","section":"Appendix C.3"},{"comment":"Reference [11] appears to misspell 'Larocca' as 'Larocza'; please verify the author name.","section":"References"},{"comment":"The relationship between Eq. (4), which uses sqrt(G_ii), and Algorithm 7, which uses the per-coordinate standard deviation of training coordinates, is not explained; the two prescriptions should be reconciled or their distinct roles stated explicitly.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is transparent about its confounds and the calibration atlas is worth publishing if reframed as a descriptive study. The abstract and conclusion overstate the validation, and I recommend major revision rather than rejection because the unsupported causal claim can be removed or qualified without changing the core calibration protocol. I would also ask the authors to provide matched-learning-rate 2-qubit runs and the full 3-qubit end-to-end evaluation before a revised submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I'll cut to it: the calibration atlas is useful and likely reproducible, but the paper's headline claim—that better-conditioned parameterizations converge faster and reach higher fidelity—is contradicted by its own matched 3-qubit comparison. That mismatch needs fixing before publication.\n\nWhat's genuinely new: a systematic study of seven density-matrix parameterizations for diffusion QST, quantified via SDR and DA from the Jacobian Gram matrix. The 2- and 3-qubit tables (Tables 2 and 3) give a clean reference for practitioners, and the per-coordinate sigma_data recipe (Section 5.2) is a sensible practical idea. The methodology is specified well enough to reproduce (finite differences, 30 states, median aggregates), and the authors are admirably transparent about protocol differences and the paradox they noticed. The boundary-effects discussion under CFG is thoughtful and does real explanatory work for the ranking reversal.\n\nThe soft spot is the central causal claim. The abstract and Section 1.3 say end-to-end training confirms that better conditioning helps. But the matched 3-qubit run (Table 8: same lr, batch, lambda_meas=0) shows Hermitian direct (2.0×) at 0.7987 versus Bloch (1.0×) at 0.4507. That's the opposite direction. At 2 qubits without CFG (Table 7), the Herm B advantage is small (+0.008 at 300 shots) and confounded by different learning rates (Herm B 4e-3, Bloch 2e-3, Appendix C). The paper honestly labels this an isometry–performance paradox and leaves the mechanism open, but that means the validation claim in the abstract is not supported. The scatter plot in Section 4.4 (Pearson −0.87, n=3) doesn't rescue it. So the strongest defensible statement is that conditioning metrics are descriptive—they map the design space—but they don't predict end-to-end ranking in a simple way, at least not without accounting for boundary effects and training details. The fix-trace recommendation is also based on calibration alone, not end-to-end training; that's fine as a heuristic but should be labeled as such.\n\nFor a referee: this paper deserves serious review. The atlas is worth having, the methodology is reproducible, and the contradictions are interesting rather than sloppy. But the abstract and Section 1.3 need revision to scoped claims. The missing code and error bars are minor by comparison. I'd send it back with a request to either run a fully matched protocol and resolve the paradox or explicitly state that conditioning metrics do not predict end-to-end performance in the absence of CFG. That would make it a solid contribution.","headline":"Useful calibration atlas, but the central claim that better conditioning predicts end-to-end performance is contradicted by the paper's own matched 3-qubit experiment.","tokens_in":12293,"tokens_out":2815,"would_cite":true,"duration_ms":23798,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper's central claim is that the Jacobian Gram matrix of a density-matrix parameterization provides a pre-training, quantitative criterion for choosing coordinate charts in diffusion-based quantum state tomography, and that without…","keywords":["quantum state tomography","diffusion models","density matrix parameterization","Jacobian Gram matrix","spectral dynamic range","diagonal anisotropy","classifier-free guidance","Cholesky factorization"],"falsifier":"Run a fully matched training comparison of Hermitian-direct and Bloch parameterizations with identical measurement-consistency loss, learning rate, and no classifier-free guidance: if Bloch reaches higher final fidelity, the claim that better-conditioned charts converge faster is falsified. A second check is to recompute the Jacobian Gram matrix on noise-perturbed samples along the diffusion trajectory and show that the fixed-state spectral dynamic range and diagonal anisotropy do not track the time-dependent conditioning the denoiser experiences.","tokens_in":11236,"feed_emoji":"⚛️","tokens_out":9237,"duration_ms":71285,"temperature":0.7,"pith_summary":"This paper tries to establish that the coordinate chart used to represent density matrices is a consequential, measurable design choice for diffusion-based quantum state tomography. It introduces the Jacobian Gram matrix $\\mathbf{J}^\\top\\mathbf{J}$ and two derived metrics, spectral dynamic range and diagonal anisotropy, as pre-training diagnostics for parameterization geometry. Calibrating seven parameterizations at two- and three-qubit scale, the paper finds that no chart is best on both isometric conditioning and physical-constraint satisfaction, and that theoretically elegant charts such as the exponential map are the worst-conditioned. End-to-end training without classifier-free guidance shows better-conditioned parameterizations converge faster and reach higher final fidelity; with guidance, the ranking reverses because guidance amplifies out-of-domain boundary effects. If correct, this gives practitioners a quantitative way to choose parameterizations and to calibrate noise scales per coordinate before expensive training.","feed_headline":"Conditioning predicts which density-matrix chart trains fastest","feed_subtitle":"A Jacobian-Gram calibration of seven charts tells diffusion QST practitioners which to pick before training.","key_machinery":"The central object is the Jacobian Gram matrix $G = \\mathbf{J}^\\top\\mathbf{J}$, where $\\mathbf{J} = \\partial\\rho/\\partial y$ is the derivative of a parameterization map $y \\mapsto \\rho$ with respect to its coordinates. From $G$ the paper extracts the spectral dynamic range, $\\kappa_{\\mathrm{spec}} = \\lambda_{\\max}(G)/\\lambda_{\\min}(G)$, and the diagonal anisotropy, $\\kappa_{\\mathrm{diag}} = \\max_i G_{ii}/\\min_i G_{ii}$, two numbers that summarize how stretched and direction-dependent the coordinate chart is. The argument runs through three steps: calibrate $G$ by finite differences on thirty sampled states, rank the seven charts by these metrics, and check the ranking against end-to-end diffusion training; the same $G$ also motivates a coordinate-aware noise-scale calibration.","core_discovery":"On the paper's own terms, the central discovery is that parameterization geometry, as measured by the Jacobian Gram matrix, is a real and largely independent axis of performance in diffusion quantum state tomography. The calibration shows that isometry and constraint enforcement are orthogonal objectives, so no single parameterization dominates both. The training experiments then show that, without classifier-free guidance, better conditioning predicts faster convergence and higher fidelity, even though the perfectly isotropic Bloch chart is overtaken by the slightly anisotropic Hermitian-direct chart; under guidance, boundary effects from out-of-domain sampling reverse the ranking. The paper ends with practical guidelines: fix-trace for the best conditioning-constraint tradeoff and per-coordinate noise-scale calibration for diffusion schedules.","pith_inferences":["The same Jacobian-Gram diagnostic could be extended to quantum process tomography, where process matrices live on a similarly constrained manifold; the paper mentions this as future scope rather than demonstrating it.","At three qubits, the calibration predicts that Hermitian-direct and Bloch remain scale-invariant, so a matched three-qubit training comparison would be a direct test of whether the two-qubit performance ordering holds.","Because guidance reverses the ranking via projection costs, deployment choices about classifier-free guidance should be made jointly with the parameterization choice rather than after it.","A testable refinement of the framework is to compute conditioning dynamically at several noise levels along the diffusion trajectory, which would reveal whether the fixed-state calibration misses time-dependent effects."],"forward_implications":["No single density-matrix parameterization can be best in both isometric conditioning and physical-constraint satisfaction, so method designers should treat the two axes as separate desiderata.","The exponential map's conditioning degrades from 149-fold at two qubits to 75,658-fold at three qubits, so by-construction physical guarantees do not imply good optimization geometry.","Without classifier-free guidance, better-conditioned parameterizations converge faster and reach higher final fidelity, so pre-training calibration can inform chart selection.","Under classifier-free guidance the ranking reverses at high shot counts because guidance amplifies out-of-domain sampling and projection costs, especially for unbounded charts.","Fix-trace is the recommended chart when automatic trace preservation is needed, and noise scales should be calibrated per coordinate rather than globally."],"supporting_citations":[{"why":"Supplies the Cholesky parameterization that diffusion QST inherits as the default baseline.","marker":"[1]"},{"why":"Establishes denoising diffusion models for quantum state tomography, whose inherited Cholesky choice this paper challenges.","marker":"[2]"},{"why":"Another diffusion-based QST method that inherits the standard parameterization and serves as prior work.","marker":"[3]"},{"why":"Supplies the diffusion training and noise-schedule machinery used in the experimental validation; its sigma-data calibration is extended to per-coordinate form.","marker":"[7]"},{"why":"Supplies classifier-free guidance, which the paper shows reverses the parameterization ranking through boundary effects.","marker":"[8]"},{"why":"Maximum-likelihood estimation baseline against which diffusion QST fidelity is compared.","marker":"[9]"}],"fun_headline_variants":["Isometry vs constraints: no single density-matrix chart dominates","Conditioning predicts diffusion QST training speed, not elegance","Jacobian Gram ranking: pick fix-trace for diffusion QST","For diffusion QST, geometric conditioning beats theoretical elegance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Jacobian Gram matrix, evaluated at a fixed set of 30 interior state samples, represents the geometry the diffusion model actually encounters when its inputs are noise-perturbed and may fall outside the valid domain.","fun_headline_variants_meta":{"raw":{"variants":["Isometry vs constraints: no single density-matrix chart dominates","Conditioning predicts diffusion QST training speed, not elegance","Jacobian Gram ranking: pick fix-trace for diffusion QST","For diffusion QST, geometric conditioning beats theoretical elegance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000691,"raw_usage":{"total_tokens":3114,"prompt_tokens":913,"completion_tokens":2201,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":2132}},"tokens_in":529,"tokens_out":2201,"duration_ms":15630,"temperature":1.0,"reasoning_tokens":2132,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:52:52.200213+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a fully matched training comparison of Hermitian-direct and Bloch parameterizations with identical measurement-consistency loss, learning rate, and no classifier-free guidance: if Bloch reaches higher final fidelity, the claim that better-conditioned charts converge faster is falsified. A second check is to recompute the Jacobian Gram matrix on noise-perturbed samples along the diffusion trajectory and show that the fixed-state spectral dynamic range and diagonal anisotropy do not track the time-dependent conditioning the denoiser experiences.","supporting_citations":[{"cited_title":"The diagonal of the multiplihedra and the tensor product of A-infinity morphisms","cited_arxiv_id":"2206.05566","evidence_quote":"Supplies the Cholesky parameterization that diffusion QST inherits as the default baseline."},{"cited_title":"Denoising diffusion models for quantum state tomography","cited_arxiv_id":null,"evidence_quote":"Establishes denoising diffusion models for quantum state tomography, whose inherited Cholesky choice this paper challenges."},{"cited_title":"Batch-in-Batch: a new adversarial training framework for initial perturbation and sample selection","cited_arxiv_id":"2406.04070","evidence_quote":"Another diffusion-based QST method that inherits the standard parameterization and serves as prior work."},{"cited_title":"Elucidating the design space of diffusion-based generative models (EDM).Advances in Neural Information Processing Systems (NeurIPS), 35, 2022","cited_arxiv_id":null,"evidence_quote":"Supplies the diffusion training and noise-schedule machinery used in the experimental validation; its sigma-data calibration is extended to per-coordinate form."},{"cited_title":"Efficient method for computing the maximum likelihood quantum state from measurements with additive Gaussian noise","cited_arxiv_id":null,"evidence_quote":"Maximum-likelihood estimation baseline against which diffusion QST fidelity is compared."}],"review_version":1}