{"id":"3b1084cd-00dc-43f5-ba0a-60313b784eb7","arxiv_id":"2508.12821","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A parallelized MPS encoding gives logarithmic-depth circuits for low-rank states, with two-CNOT gates and grid-topology adaptation, validated numerically on smooth functions and distributions.","lead":"This paper proposes a parallelized matrix-product-state method that loads smooth data into quantum states with far fewer time steps, and it adapts the circuit to grid-shaped hardware. The practical payoff is that current noisy quantum computers could run finance and simulation algorithms with lower error.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The O(log n) preparation claim lacks a fidelity bound for the chi=2 truncation; for non-contiguous TTN/HTN pairings the truncated matrix can have rank up to chi^2, so exact rank-2 preparation is unproven.","rationale":"The reader identified the chi=2 truncation as the weakest assumption, which is the right root cause. My concern sharpens this: the relevant object is not the global Schmidt rank but the rank of the 4 x 2^(n-2) matrix A_{a,b} for the specific (possibly non-contiguous) pairs used in the parallel schedules. That rank can exceed chi for legitimate rank-2 MPS states, so the exactness claim in Sec. 2.3 is not justified by the MPS bond-dimension bound, and the depth-only argument in Proposition 1 does not guarantee approximation quality. This is a correctness risk to the central claim, not merely a missing detail. A concrete random-rank-2 MPS test would settle whether the exactness claim is true for the proposed schedules. If it fails, the O(log n) result would need to be restated as an uncontrolled approximation, or restricted to states for which every scheduled gate's A_{a,b} has rank at most 2. I do not think this changes the overall verdict from CONDITIONAL, since the numerical evidence on the tested functions is positive and the missing proof of Theorem 1 is a separate, less central issue; hence UNCHANGED.","tokens_in":9509,"tokens_out":16548,"duration_ms":176263,"concrete_test":"Generate a random exact MPS with bond dimension chi=2 on n=8 qubits (random complex tensors, normalized), and construct the IMPS circuit exactly as described for the TTN/HTN schedule, keeping only the two largest singular values at each SVD step and then inverting the circuit to prepare the state. Compute the fidelity with the target state and repeat for, say, 20 random samples. If the fidelity is not 1 (or not bounded below by 1 - 1e-6) for generic samples, the claim that all MPS rank-2 states are precisely prepared in O(log n) depth is falsified; if it is 1, the exactness claim survives for n=8 and would need a rigorous proof for general n.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central scaling claim rests on the assertion that keeping the two largest singular values at each SVD step (Eq. 2-3) gives a good approximation, and Sec. 2.3 further claims that states with MPS rank 2 can be precisely prepared in O(log n) depth. For a rank-2 MPS, any contiguous bipartition has Schmidt rank at most 2, so A_{a,b} has rank at most 2 when the gate pair (a,b) is separated by a single bond. However, the TTN and HTN schedules pair qubits that are not adjacent in the original order (e.g., HTN first pairs i and i+4 after binary reordering; TTN pairs (2,3), whose block has both left and right boundaries). For such pairs, the rank of A_{a,b} is bounded by the product of the two boundary bond dimensions, i.e., up to chi^2 = 4, not chi = 2. Truncating to two singular values then discards nonzero weight even for exactly representable chi=2 states. The paper gives no cumulative error bound over the O(log n) layers, and Proposition 1 is argued only from the number of parallel gates, not from fidelity. The numerical comparisons also use different truncation protocols for HTN/HEN versus MPS/TTN, as the paper itself states in Sec. 3, so the reported fidelity advantage is not a controlled comparison at fixed approximation error. Thus the load-bearing assumption that bounded global MPS rank controls the per-gate truncation error is not established and is false for generic non-contiguous pairings.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes an improved Matrix Product State (IMPS) amplitude-encoding protocol. The method applies SVD-based two-qubit disentangling gates in parallel at each stage, which the authors claim yields O(log n) circuit depth for states with bounded MPS rank. A second ingredient uses a Cartan/KAK decomposition to replace each generic two-qubit unitary by an element of an equivalence class implementable with two CNOT gates instead of three. The paper defines a ring of bounded-MPS-rank functions and reports numerical comparisons on three functions and three distributions using MPS, TTN, HEN, and HTN connectivity schemes, claiming higher fidelity and lower depth, together with a discussion of grid-topology adaptation.","tokens_in":9845,"tokens_out":15316,"duration_ms":161065,"significance":"If established, the claimed log-depth preparation of bounded-MPS-rank amplitude states with roughly one-third fewer two-qubit gates would be practically useful for NISQ-era amplitude encoding and for quantum Monte Carlo integration. The SVD-based parallel disentangling idea is transparent, the equivalence-class construction around Eq. (6) is a useful observation, and the numerical experiments cover relevant financial and distribution-type benchmarks. However, the central proposition and theorem are not supported by the submitted text: Proposition 1 has no error analysis, and Theorem 1's proof is in a missing appendix. In addition, the numerical comparison changes the truncation/reset protocol between method families, and the benchmarks are generated from the same bounded-rank ring that the method assumes. These gaps currently limit the significance of the results.","major_comments":[{"comment":"Proposition 1 is asserted with a gate-counting argument rather than a proof. The observation that floor(n/2) two-qubit gates can be applied in parallel per layer only establishes a circuit-depth upper bound if the truncations are exact or have controlled errors. No cumulative error bound for O(log n) layers of chi=2 truncation is provided, and no theorem connects the global MPS rank of the target state to the error of the intermediate SVDs. The proposition should be replaced by a precise statement with proof for the exact rank-2 case and a separate approximate statement with an error bound in terms of the discarded singular values.","section":"Sec. 2.1, Proposition 1"},{"comment":"The proof of Theorem 1 is said to be provided in Appendix A, but no Appendix A appears in the manuscript. The two-CNOT decomposition is one of the paper's main contributions, so this proof must be included. Moreover, the theorem as worded ('Any fourth-order unitary matrix arising from the SVD in MPS amplitude preparation') is too broad: for a 4 by 2^{n-2} amplitude matrix the left singular matrix can be an arbitrary 4x4 unitary, and a generic two-qubit unitary requires three CNOT gates. The theorem should be restricted to the equivalence-class matrices constructed in Eq. (6), with a proof that those matrices admit a two-CNOT decomposition.","section":"Sec. 2.2, Theorem 1"},{"comment":"The fidelity comparison is not controlled. Section 3 states that for MPS and TTN 'truncation involves resetting the amplitude to its original length after each layer,' whereas for HTN and HEN 'the amplitude length is reset after each U-depth.' Because the reported fidelity advantage of HTN/HEN is the paper's main empirical claim, the comparison must use identical reset/truncation rules for all four methods, or both rules must be reported for each method. In addition, Fig. 8 compares the chip-adapted IMPS at U-depths 5 and 10 with the traditional MPS at U-depth 11, so the claim of 'comprehensively surpasses its fidelity' should be stated with the depth difference made explicit.","section":"Sec. 3, numerical comparisons"},{"comment":"The statement that 'we can precisely prepare states with an MPS rank of 2' in O(log n) depth is not proved for the schedules actually used. For a non-contiguous pair (a,b), the matrix A_{a,b} in Eq. (1) has rank bounded by the product of the two boundary bond dimensions, which is up to chi^2 = 4 even for a rank-2 MPS, so truncating to the two largest singular values is not exact. The exactness claim appears plausible for the TTN schedule, where each pair is adjacent in the residual order at each stage, but it must be stated for that schedule and proved, and the HTN/HEN schedules need a separate error analysis.","section":"Sec. 2.3, exact rank-2 claim"}],"minor_comments":[{"comment":"The terms U-depth and layer are defined at the end of the introduction, but in Section 3 'layer' is used for both the core preparation block and for a sequential two-qubit gate layer; please use one consistent terminology.","section":"Sec. 1"},{"comment":"The reference list contains a duplicated entry: references [46, 47, 47] are cited in the sentence on reinforcement learning; please correct.","section":"Sec. 2.1"},{"comment":"The 'well-known theorem for tensor networks' on the ring structure is stated without a citation or proof; please add a reference or a proof sketch.","section":"Sec. 2.3, Theorem 2"},{"comment":"The caption lists the disentangled qubits but not the gate pairs for each layer; for reproducibility, list the pairs acted on in each layer (e.g., second layer pairs (1,3) and (5,7)).","section":"Fig. 2"},{"comment":"The numerical benchmarks are all generated from the bounded-rank ring structure introduced in Sec. 2.3; while this is consistent with the stated scope, a test on states with known MPS rank outside this construction would strengthen the claimed broad applicability.","section":"Sec. 3"}],"recommendation":"major_revision","confidential_remarks":"The paper should be positioned more carefully against Ref. [37], which already proves log-depth preparation of MPS states; the incremental contribution here is the specific SVD-based parallel construction and the CNOT reduction. The missing Appendix A and the uncontrolled numerical comparison make a decision difficult at this stage, and I recommend re-review after these are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the IMPS paper. Here's the short version: the paper has one useful engineering idea and a mostly known theoretical claim. The HTN mapping to grid chips is a concrete scheme that could be handy for practitioners, and the two-CNOT decomposition of the SVD unitaries is a nice constant-factor optimization that they verify numerically. But the O(log n) depth claim for MPS states is already in Malz et al. [37], which they cite; Proposition 1 is an assertion, not a proof. The real problem is that the paper does not establish a fidelity bound for the SVD truncation, and the stress-test note about non-contiguous pairings is correct: for a generic rank-2 MPS, the amplitude matrix for a non-adjacent qubit pair can have rank up to 4, not 2, so truncating to two singular values can discard weight. The paper's own benchmark functions (exponentials, cosines, and their products) are sums of two product states, which are special, so the numerics survive, but the general claim 'precise preparation for MPS rank 2' is too strong.\n\nOther soft spots: the fidelity comparison between HTN/HEN and MPS/TTN uses different truncation protocols (per layer vs per U-depth), so the reported fidelity advantage is not a controlled comparison; Theorem 1's proof is in a missing Appendix A. The paper also overstates the depth reduction: the depth drops from O(n) to O(log n), but the total gate count is still O(n) up to the constant factor saved by the two-CNOT trick.\n\nWhat's genuinely useful: the HTN construction with depth O((m+n)/2) on an n x m grid (O(sqrt(N)) for square grids) is an explicit recipe. The two-CNOT gate form is a special case of KAK, but a good one to have written down. The numerics on the six test functions are clean and support the method on that set.\n\nThis paper deserves a serious referee, not a desk reject. A referee should ask for the missing appendix, a clear comparison with [37], and a controlled fidelity benchmark. If the authors supply those, the HTN mapping and gate optimization are worth publishing. As it stands, I'd treat the scaling claim as heuristic.\n\nTake it to reading group if you want a case study in how a useful engineering paper can overreach its theory. I wouldn't cite it in my own work yet.","headline":"Useful topology-aware MPS preparation recipe, but the log-depth claim is known and the fidelity argument is weaker than advertised.","tokens_in":10387,"tokens_out":8044,"would_cite":false,"duration_ms":85345,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Improved MPS preparation loads bounded-rank amplitude states in $O(\\log n)$ depth and needs only two CNOTs per two-qubit gate.","keywords":["quantum state preparation","amplitude encoding","matrix product state","logarithmic circuit depth","two-qubit gate decomposition","CNOT count","tensor network","bounded MPS rank"],"falsifier":"Take an $n$-qubit state whose Schmidt decomposition across some middle cut has three equal nonzero coefficients (bond dimension 3). IMPS with $\\chi=2$ truncates one of the three Schmidt modes at that cut, so the overlap of the prepared state with the exact target is at most $2/3$ there; computing this overlap for increasing $n$ will show the logarithmic-depth circuit does not prepare the exact state and that fidelity saturates below one, isolating the bound on the claim.","tokens_in":9305,"feed_emoji":"⚛️","tokens_out":7756,"duration_ms":79653,"temperature":0.7,"pith_summary":"The paper aims to show that amplitude-encoded quantum states with bounded matrix-product-state rank can be prepared by a unitary circuit of logarithmic depth, instead of the linear depth of standard MPS methods. The key move is to apply SVD-based two-qubit disentanglers in parallel layers, approximately zeroing one qubit per gate so the number of entangled qubits halves each round. A second claim is that every two-qubit unitary arising from the SVD can be implemented with two CNOT gates plus single-qubit gates, cutting the two-qubit gate count by roughly one third for complex amplitudes. If correct, this makes loading smooth functions and financial distributions practical on near-term hardware, including grid-connected chips where the depth scales as the square root of the qubit count.","feed_headline":"Log-depth amplitude loading with a third fewer CNOTs","feed_subtitle":"Improved matrix-product-state method cuts encoding depth from linear to logarithmic and maps onto grid chips.","key_machinery":"The carrying object is the SVD of the $4\\times 2^{n-2}$ amplitude matrix $A_{a,b}$: applying $U^{-1}$ from $A=USV^\\dagger$ converts the amplitudes into diagonal singular-value rows, so one qubit is disentangled up to truncation. This makes two-qubit gates position-independent, enabling parallel layers of $\\lfloor n/2\\rfloor$ gates that halve the number of active qubits per round (the tree/hypercube tensor-network ordering) and yield $O(\\log n)$ depth. The second mechanism is the unitary equivalence class $\\{\\mathrm{diag}(U_1,U_2)U^{-1}\\}$, combined with the cosine-sine/KAK decomposition, which produces the two-CNOT implementation of each two-qubit gate. A third supporting piece is the ring subadditivity/submultiplicativity of MPS ranks, which certifies that sums and products of basic function states remain preparable at bounded cost.","core_discovery":"Starting from an MPS with bond dimension $\\chi=2$, the protocol builds a $4\\times 2^{n-2}$ amplitude matrix $A_{a,b}$ for any pair of qubits $(a,b)$, performs the SVD $A=USV^\\dagger$, and applies $U^{-1}$; the state of qubit $a$ becomes the two dominant singular-value rows, so after truncation that qubit is approximately $|0\\rangle$. Because the pair is arbitrary, the paper pairs disjoint qubits into a single parallel layer, disentangling roughly half the remaining qubits per layer, which gives the claimed $O(\\log n)$ depth via a tree- or hypercube-tensor-network ordering. For two-qubit unitaries, the paper uses the freedom to left-multiply by an arbitrary block-diagonal unitary and a cosine-sine (KAK) decomposition to obtain a representative implemented by two CNOT gates and single-qubit rotations, stated as Theorem 1. The paper further proves subadditivity and submultiplicativity of MPS rank for element-wise sums and products of function states, identifying a ring of bounded-rank functions that includes cosine, linear, exponential-derived, and common distribution states, and gives numerical evidence that the hypercube ordering beats conventional MPS and TTN in fidelity at matched depth.","pith_inferences":["An implicit consequence is that the equivalence-class trick behind the two-CNOT decomposition is not tied to MPS: any SVD-based two-qubit synthesis with the same block-unitary and cosine-sine structure could inherit the one-third CNOT saving, so it is worth testing on generic unitary compilation pipelines.","The paper's edge-contraction view turns chip mapping into a graph problem; a testable extension is to search contraction orders that minimize truncation error, not just depth, which could improve fidelity on irregular topologies beyond the grid example.","The rank-ring result suggests a generative recipe: if a financial payoff or probability density can be written as a short sum or product of exponentials, cosines, and polynomials, its MPS rank is bounded by a simple formula, making efficient loading a syntactic check rather than a numerical search."],"forward_implications":["Bounded-rank states such as cosine, linear, GHZ, and W states—and any element-wise sum or product built from exponentials, cosines, and linear functions—are preparable exactly in $O(\\log n)$ depth.","Complex-valued amplitude loading uses roughly one third fewer CNOT gates per two-qubit unitary, directly cutting the dominant error source on noisy hardware.","On an $n\\times m$ grid chip, the same construction prepares a state in depth about $(m+n)/2$, i.e. $O(\\sqrt{N})$ for a square grid, and rank-2 function states in $O(\\max(m,n))$.","The protocol remains unitary and does not require mid-circuit measurement or ancillas, so the prepared states can be reused inside amplitude estimation or quantum Monte Carlo subroutines."],"supporting_citations":[{"why":"Supplies the prior log-depth MPS preparation result that this work makes practical and unitary.","marker":"[37]"},{"why":"Provides the constant-depth adaptive-circuit result that frames the unitary-depth optimum context.","marker":"[38]"},{"why":"Gives the conventional MPS-to-quantum-circuit mapping whose linear depth is the baseline the paper improves upon.","marker":"[29]"},{"why":"Provides the normal-distribution MPS preparation method used as a fidelity and depth benchmark.","marker":"[35]"},{"why":"Gives the three-CNOT universal two-qubit decomposition whose gate count Theorem 1 reduces to two.","marker":"[43]"},{"why":"Supplies the Cartan/KAK decomposition used in the two-qubit gate optimization.","marker":"[44]"},{"why":"Justifies that an MPS with sufficiently large bond dimension can represent any quantum state, grounding the approximation.","marker":"[45]"}],"fun_headline_variants":["Log-depth encoding via improved MPS, 33% fewer CNOTs","Exponential circuit-depth cut for amplitude loading","MPS protocol drops depth to log scale, cuts gates","Grid-friendly MPS encodes states at log depth"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole depth and fidelity guarantee rests on truncating every SVD to its two largest singular values (bond dimension $\\chi=2$); if the target state has substantial weight in the discarded singular vectors, the logarithmic-depth circuit only produces an approximation whose fidelity can be far from one.","fun_headline_variants_meta":{"raw":{"variants":["Log-depth encoding via improved MPS, 33% fewer CNOTs","Exponential circuit-depth cut for amplitude loading","MPS protocol drops depth to log scale, cuts gates","Grid-friendly MPS encodes states at log depth"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000631,"raw_usage":{"total_tokens":2900,"prompt_tokens":917,"completion_tokens":1983,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":1917}},"tokens_in":533,"tokens_out":1983,"duration_ms":15185,"temperature":1.0,"reasoning_tokens":1917,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:20:23.111930+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take an $n$-qubit state whose Schmidt decomposition across some middle cut has three equal nonzero coefficients (bond dimension 3). IMPS with $\\chi=2$ truncates one of the three Schmidt modes at that cut, so the overlap of the prepared state with the exact target is at most $2/3$ there; computing this overlap for increasing $n$ will show the logarithmic-depth circuit does not prepare the exact state and that fidelity saturates below one, isolating the bound on the claim.","supporting_citations":[{"cited_title":"Preparation of matrix product states with log-depth quantum circuits.Physical Review Letters, 132(4):040404, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the prior log-depth MPS preparation result that this work makes practical and unitary."},{"cited_title":"Constant-depth preparation of matrix product states with adaptive quantum circuits","cited_arxiv_id":null,"evidence_quote":"Provides the constant-depth adaptive-circuit result that frames the unitary-depth optimum context."},{"cited_title":"Encoding of matrix product states into quantum circuits of one-and two-qubit gates.Physical Review A, 101(3):032310, 2020","cited_arxiv_id":null,"evidence_quote":"Gives the conventional MPS-to-quantum-circuit mapping whose linear depth is the baseline the paper improves upon."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies that an MPS with sufficiently large bond dimension can represent any quantum state, grounding the approximation."}],"review_version":2}