{"id":"e7ade810-7a32-4baa-86fb-6dd3df57684a","arxiv_id":"2601.09037","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A fully on-chip FPGA p-bit parallel-tempering solver with graph sparsification detects 64x64 BPSK MIMO below MMSE at 4.7 ms/instance, and 2D parallel tempering removes the copy-constraint tuning bottleneck with ~250-500x faster SK convergence.","lead":"This paper builds a probabilistic computer on an FPGA that detects wireless signals by solving a hard optimization problem, and it introduces a two-dimensional temperature/constraint-swapping algorithm to remove a manual tuning step. A generalist should care because it is a concrete, fully on-chip demonstration of a p-bit Ising machine on a real dense problem, with measured latency and projected ASIC performance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract claims on-chip 2D-PT validation on 16x16 MIMO, but Section V states FPGA implements only 1D-PT; the second on-chip solver is unsupported.","rationale":"I read the full manuscript and the reader's verdict. The reader's weakest_assumption focused on sparsification and P-sensitivity, which is a legitimate concern but not the most load-bearing issue. A more direct problem is the contradiction between the abstract and the body regarding on-chip 2D-PT. The abstract claims a 16×16 MIMO on-chip validation, but Section V explicitly says the FPGA only implements 1D-PT. The body contains no 16×16 MIMO experiment, no 54-replica configuration, and no on-chip 2D-PT results. The 2D-PT speedup claims are therefore simulated only, contrary to the abstract's implication of a second fully on-chip solver. This directly undermines the paper's central advertised contribution. While I do not question the authors' integrity, the inconsistency must be resolved before the claims can be taken at face value. The reader's CONDITIONAL verdict remains appropriate, but the conditions should include addressing this contradiction alongside the numerical discrepancies. I therefore set verdict_should_be to UNCHANGED (the reader's category is preserved) while disagreeing with the reader's identification of the weakest assumption, since my concern is different and more fundamental.","tokens_in":12587,"tokens_out":7131,"duration_ms":64863,"concrete_test":"Search the full text for any description of the 16×16 MIMO on-chip 2D-PT run, including the 54-replica grid, scheduling, BER data, and FPGA resource usage. If no such description exists, request that the authors either (1) provide the missing experimental details, or (2) revise the abstract to state that 2D-PT was validated only in simulation and that only 1D-PT was implemented on-chip. Also request reconciliation of the abstract/body numerical discrepancies (1,408 vs 1,920 p-bits, 3 vs 4.7 ms, 250× vs 500×) as part of the same revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's advertised central claim is 'two fully on-chip parallel tempering solvers.' The abstract explicitly states: 'We further validate 2D-PT entirely on-chip with 54 replicas (1,728 p-bits) on a 16 × 16 MIMO instance...' However, Section V's final paragraph says the opposite: 'our current FPGA only implements 1D-PT due to capacity limits, its modular architecture ... is designed to support 2D-PT in future ASIC implementations.' No 16×16 MIMO experiment, 54-replica configuration, or on-chip 2D-PT result appears anywhere in the body or Methods. The 2D-PT results for SK and 128×128 MIMO are explicitly described as Matlab simulations. This is not a secondary P-sensitivity caveat; it is a direct contradiction about whether the key algorithmic contribution was ever validated in hardware. If the body is accurate, the abstract overstates the contributions, and the paper should be revised to clarify that 2D-PT is simulated only. The reader's verdict missed this contradiction and instead assumed the 16×16 on-chip run exists. The numerical inconsistencies (1,408 vs 1,920 p-bits, 3 vs 4.7 ms, 250× vs 500×) are consistent with version drift, but the missing 2D-PT hardware is the load-bearing issue.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes graph sparsification with auxiliary copy variables to make dense optimization problems (MIMO detection, SK spin glasses) tractable for p-bit computers. It reports an FPGA implementation of 1D parallel tempering with 15 replicas of a 128-node sparsified 64×64 MIMO problem (1,920 p-bits) and claims complete end-to-end solution times of 4.7 ms per instance, with BER below MMSE. It also introduces 2D parallel tempering (2D-PT), which exchanges replicas along both temperature and copy-constraint dimensions, and reports simulated speedups over tuned 1D-PT on SK and 128×128 MIMO, plus an abstract-level claim of an on-chip 16×16 MIMO 2D-PT validation. The body states that the FPGA implements only 1D-PT and that 2D-PT results are from Matlab simulation; no on-chip 2D-PT experiment appears in the manuscript.","tokens_in":12831,"tokens_out":6493,"duration_ms":62118,"significance":"If fully substantiated, the fully on-chip 1D-PT FPGA demonstration is a useful engineering result: it shows that sparsification plus on-chip replica exchange can handle dense MIMO instances beyond the size typical of dense p-bit implementations, with claimed end-to-end overhead accounting. The 2D-PT method is conceptually appealing as a way to avoid manual penalty-strength tuning, and the simulated speedups on SK and MIMO are potentially interesting. However, the paper's advertised hardware validation of 2D-PT is contradicted by the body, and the central quantitative claims contain internal inconsistencies. The manuscript also does not provide publicly available code or data, relying on 'available upon request.' The contribution therefore cannot be assessed at its face value until the scope and numbers are corrected.","major_comments":[{"comment":"The abstract claims 'We further validate 2D-PT entirely on-chip with 54 replicas (1,728 p-bits) on a 16 × 16 MIMO instance' and describes 'two fully on-chip parallel tempering solvers.' Section V explicitly states: 'our current FPGA only implements 1D-PT due to capacity limits' and that the SK and 128×128 MIMO 2D-PT experiments are performed in Matlab simulation. No 16×16 MIMO, 54-replica, or on-chip 2D-PT result appears in the body, figures, or Methods. This is not a minor presentation issue: the paper's advertised central hardware claim is unsupported. The abstract, conclusion, and any title-level framing must be revised to state that 2D-PT is validated in simulation only.","section":"Abstract vs. Section V, final paragraph"},{"comment":"The paper claims 2D-PT is 'tuning-free' or removes manual parameter optimization, but the adaptive schedule uses hyperparameters αβ and αP that are 'initially tuned on a single instance,' and the final β and P schedules are averaged over 10 instances. No sensitivity analysis is provided for αβ, αP, or the stopping thresholds. Without such analysis, the reported 500× speedup and 10× MIMO acceleration could reflect the hand-selected schedules rather than the method itself. A concrete test would be to vary αβ and αP over a reasonable range and report the distribution of residual energy/BER, or to fix the parameters across problem sizes and instance distributions.","section":"Section V, Methods B, Eq. (12)"},{"comment":"The quantitative claims are mutually inconsistent across the front matter and body. The front-matter abstract cites 1,408 p-bits and 3 ms/instance; the body says 1,920 p-bits and 4.7 ms/instance. The abstract reports ~250× faster 2D-PT convergence; Section V reports 500×. The body also says timing is averaged over 13,000 instances, while the BER results are averaged over 100 channels × 10 transmitted vectors = 1,000 instances. These numbers are headline results and must be reconciled or clearly labeled as different configurations/versions.","section":"Abstract, Section III, Section V"},{"comment":"For the 128×128 MIMO 2D-PT result, BER = 0 at high SNR is reported for 200 channels and 10 transmitted symbols per channel, which is 256,000 bits. On a finite test set, BER = 0 only provides an upper bound (roughly < 1.2×10^-5 at 95% confidence). The paper should state this bound and, ideally, report the number of bit errors or the confidence interval rather than a literal zero, especially when comparing with 1D-PT's error floor near 10^-5.","section":"Section V, Fig. 5a"}],"minor_comments":[{"comment":"The caption states 'all timing values are averaged over 13,000 instances,' but the text says results are averaged over 100 channel realizations each decoding 10 transmitted vectors, i.e., 1,000 instances. Please correct the count or clarify what was measured.","section":"Section III, Fig. 2"},{"comment":"The abstract mentions ASIC projections of 103 MHz and 285.8 mW for the 16×16 2D-PT design, while Section IV gives 89 MHz and 185.2 mW for the 64×64 1D-PT design. The relationship between these projections is not discussed in the body, and the 16×16 ASIC results are absent. Clarify which design each projection refers to.","section":"Section IV, Abstract"},{"comment":"The table reports β and P schedules with inconsistent decimal precision (e.g., 0.500 vs. 0.801 vs. 24.6). Use a consistent format and indicate the number of significant digits. Also note that the MIMO P column has 13 entries, matching the 13 columns, but the β column has 16 entries; this should be explicitly stated in the caption.","section":"Section V, Table II"},{"comment":"The base-2 exponential approximation uses a floor operation in the exponent but the subsequent 'fractional multiplication factor 23/16' description is unclear. Please define the approximation precisely so that the fixed-point implementation can be reproduced.","section":"Methods A, Eq. (9)"},{"comment":"The text says 2D-PT finds 'constrained ground states (ρ=0)' in fewer than 1,000 swaps. The figure shows residual energy down to 10^-6, not exactly zero. Specify whether ρ=0 is exact, within numerical tolerance, or a statistical statement.","section":"Section V, Fig. 4c"}],"recommendation":"major_revision","confidential_remarks":"The abstract/body mismatch on on-chip 2D-PT is serious enough that the manuscript as submitted cannot be evaluated for its advertised claims. The 1D-PT FPGA demonstration and the simulated 2D-PT results are both potentially publishable contributions, but the authors must first decide on a single, consistent set of hardware claims and fix the numerical discrepancies. I would ask the editor to require a revised version that clearly separates hardware-validated results from simulations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things before you read it. The on-chip 1D-PT result is real: 15 replicas of a 128-node sparsified 64x64 MIMO (1,920 p-bits) run entirely on an FPGA, with measured end-to-end 4.7 ms/instance including all overhead, BER clearly below MMSE, and a plausible 7nm ASIC projection at 89 MHz/185 mW. That is a legitimate engineering contribution. The second thing is that the abstract overclaims: it says 2D-PT is \"validated entirely on-chip\" on 16x16 MIMO with 54 replicas, but Section V explicitly says the current FPGA only implements 1D-PT and 2D-PT is left for future ASIC designs. No 16x16 experiment appears anywhere in the full text. I re-read the relevant paragraphs; this is a direct contradiction, not a nuance. The reader's report missed it.\n\nWhat the paper does well: the timing model is explicit (load, sweep, energy, swap, read, verify), the FPGA method section is detailed (fixed-point, graph coloring, swap controller), and the 1D-PT hardware claims are reproducible in principle. The application of the group's earlier 2D-PT algorithm (Ref. 26) to MIMO and SK is a useful extension, and the P-sensitivity plots in Fig. 4a are honest about the tuning problem.\n\nWhere it is soft: (1) The abstract and body disagree on almost every headline number: 1,408 vs 1,920 p-bits, 3 vs 4.7 ms, 103 vs 89 MHz, 250x vs 500x. Looks like version drift, but it must be fixed. (2) The 2D-PT speedup is simulated only; the \"tuning-free\" claim is not demonstrated in hardware. (3) The adaptive schedules use αβ/αP initially tuned on a single instance, then averaged over 10; that is a tuning step, just at a different level. It is not fatal, but it should be labeled. (4) The sparsification landscape concern is real — Fig. 4a shows sharp P sensitivity for 1D-PT — but the paper acknowledges this and 2D-PT is a sensible response, even if only simulated.\n\nBottom line: this paper deserves a serious referee, but the abstract must be revised to match the body. The 1D-PT hardware result is solid enough to publish after a careful reconciliation; the 2D-PT claims should be clearly marked as simulation. I'd send it to review, not desk reject.","headline":"Solid on-chip 1D-PT hardware result for 64x64 MIMO, but the abstract overclaims on-chip 2D-PT validation that the body explicitly says does not exist; needs a revision, not a desk reject.","tokens_in":13523,"tokens_out":3064,"would_cite":true,"duration_ms":27861,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sparsified p-bit computers with 2D parallel tempering can decode large MIMO channels near-optimally without manual tuning.","keywords":["probabilistic computing","p-bits","MIMO detection","graph sparsification","parallel tempering","Ising machines","FPGA","combinatorial optimization"],"falsifier":"On a 256×256 BPSK MIMO instance with two copies per node, run 2D-PT with the adaptive schedule from the paper; if the best decoded BER at 20 dB SNR does not go below the MMSE baseline within 10^5 swaps, or if the copy-agreement percentage in the converged replicas stays below 95%, the claim that sparsification preserves the global optimum landscape is falsified.","tokens_in":12327,"feed_emoji":"📡","tokens_out":8641,"duration_ms":68550,"temperature":0.7,"pith_summary":"Probabilistic computers built from stochastic p-bits can solve dense combinatorial problems, but the all-to-all connectivity required by real-world instances like MIMO detection has scaled poorly in hardware. This paper argues that graph sparsification with auxiliary copy variables, plus a novel two-dimensional parallel tempering (2D-PT) algorithm that exchanges replicas across both temperature and copy-constraint strength, solves that scaling problem. The authors demonstrate a fully on-chip FPGA solver with 1,920 p-bits that decodes 64×64 BPSK MIMO with bit error rates well below the MMSE baseline in 4.7 ms end-to-end, and show in simulation that 2D-PT converges about 500× faster than optimally tuned 1D parallel tempering on Sherrington-Kirkpatrick spin glasses and reaches zero bit errors on 128×128 MIMO where 1D-PT has an error floor. The central value is that the tuning bottleneck of sparsified Ising machines — the ferromagnetic copy-coupling strength P — is eliminated by the second dimension, making dense optimization practical on sparse hardware.","feed_headline":"2D parallel tempering removes tuning for sparse p-bit MIMO decoders","feed_subtitle":"On-chip p-bit solver decodes 64x64 MIMO in 4.7 ms; 2D-PT reaches zero errors on 128x128 channels.","key_machinery":"The central mechanism is the p-bit update rule combined with graph sparsification: every original node gets a small number of auxiliary 'copy' nodes connected by ferromagnetic couplings of strength P, converting the all-to-all Hamiltonian into a bounded-degree graph. The second key piece is 2D-PT, which places replicas on a two-dimensional grid indexed by inverse temperature β and copy strength P, and performs Metropolis exchanges along both axes. The P-axis swaps let feasible low-energy states from soft-constraint replicas migrate into hard-constraint replicas, enforcing copy agreement while maintaining fast mixing. The FPGA implementation uses 10-bit fixed-point arithmetic, a Taylor approx","core_discovery":"The paper establishes that dense MIMO detection can be mapped to a sparse p-bit network using auxiliary copy variables, and that a fully on-chip parallel tempering solver on an FPGA — 15 temperature replicas, 1,920 p-bits — decodes 64×64 BPSK MIMO with bit error rates substantially below the MMSE detector, with complete end-to-end solution times of 4.7 ms per instance including all overheads. It then shows that the main weakness of sparsification, the sharp sensitivity to the copy-constraint strength P, is removed by 2D-PT, which swaps replicas along both temperature and P. In simulation, 2D-PT converges about 500× faster than optimally tuned 1D-PT on 64-spin Sherrington–Kirkpatrick spin gla","pith_inferences":["The 2D-PT idea generalizes beyond sparsification: any optimization with a tunable penalty (e.g., a Lagrange multiplier or regularization strength) could use a second replica axis to anneal both temperature and penalty, making it a general hyperparameter-free annealing scheme.","A careful reader should note that the abstract's on-chip 2D-PT validation for 16×16 MIMO is not described in the body's methods or results; the body states the FPGA implements 1D-PT only, so the hardware evidence for 2D-PT currently rests on the simulation results unless that discrepancy is resolved.","A natural testable extension is to apply the same sparsification-plus-2D-PT pipeline to non-BPSK MIMO (e.g., QAM) and to channel sizes beyond 128×128, where the copy-variable overhead and P-sensitivity may grow; the paper's own Fig. 4a suggests performance degrades sharply outside an optimal P window, so scaling behavior is an open question.","The residual-energy versus swap-budget scaling (power-law at high SNR) suggests the solver's throughput-versus-accuracy trade-off can be tuned per SNR; a practical decoder could choose the number of steps adaptively, and 2D-PT's faster convergence directly buys latency."],"forward_implications":["If sparsification plus 2D-PT holds, dense Ising problems no longer require dense hardware; sparse chips with fixed-degree routing can scale to thousands of p-bits at constant clock frequency.","MIMO detection on 64×64 and 128×128 BPSK channels can be solved near-optimally (BER below MMSE, zero at high SNR) with end-to-end latencies of milliseconds on current FPGAs, and faster on ASICs.","The elimination of manual P-tuning makes sparsified Ising machines practical as black-box solvers for other constrained optimization problems, since the 2D-PT schedule is adaptive.","Parallel arrays of low-power ASICs (projected around 90–110 MHz and 124–185 mW in 7 nm) could approach the throughput demands of next-generation wireless receivers.","The 2D-PT framework applies to any Hamiltonian with a tunable penalty, not just copy constraints in MIMO, potentially generalizing the approach across constrained NP-hard problems."],"fun_headline_variants":["2D parallel tempering kills tuning in p-bit MIMO decoding","On-chip p-bits decode MIMO with zero errors via 2D tempering","Sparse p-bit MIMO solver gets 500x faster with 2D-PT","P-bit MIMO: no tuning, 2D-PT hits zero errors on 128x128"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that sparsifying the all-to-all graph with two auxiliary copies per node and a moderate ferromagnetic coupling P preserves the global optimum landscape closely enough that low-energy sparse states project back to near-optimal dense solutions; if this mapping is biased for larger or denser problems, the BER claims collapse.","fun_headline_variants_meta":{"raw":{"variants":["2D parallel tempering kills tuning in p-bit MIMO decoding","On-chip p-bits decode MIMO with zero errors via 2D tempering","Sparse p-bit MIMO solver gets 500x faster with 2D-PT","P-bit MIMO: no tuning, 2D-PT hits zero errors on 128x128"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000209,"raw_usage":{"total_tokens":1361,"prompt_tokens":979,"completion_tokens":382,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":723,"completion_tokens_details":{"reasoning_tokens":290}},"tokens_in":723,"tokens_out":382,"duration_ms":3980,"temperature":1.0,"reasoning_tokens":290,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T10:42:09.800270+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a 256×256 BPSK MIMO instance with two copies per node, run 2D-PT with the adaptive schedule from the paper; if the best decoded BER at 20 dB SNR does not go below the MMSE baseline within 10^5 swaps, or if the copy-agreement percentage in the converged replicas stays below 95%, the claim that sparsification preserves the global optimum landscape is falsified.","supporting_citations":[],"review_version":1}