REVIEW 4 major objections 5 minor 11 references
A 354Mb/s 0.37mm^2 151mW 32-User 256-QAM Near-MAP Soft-Input Soft-Output Massive MU-MIMO Data Detector in 28nm CMOS
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A 0.37 mm² ASIC delivers near-MAP detection for 32-user massive MIMO.
desk verdict A real chip with real measured numbers, but the near-MAP claim runs ahead of the evidence: no MAP baseline and no 256-QAM symmetric error-rate plot. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is LAMA, large-MIMO approximate message passing, an iterative algorithm that decouples the MIMO channel into parallel AWGN channels and denoises them with posterior mean and variance updates plus an Onsager correction. The chip modifies LAMA by operating on the U×U Gram matrix instead of the full channel matrix, by damping the variance and SINR updates to stabilize finite-dimensional and correlated channels, and by computing the message mean and variance in the bit domain using the max-log approximation, which approximates a sum of exponentials by its largest term. The 32×32 matrix-vector multiplication is carried out with a fan-out-reducing arrangement based on Cannon's algorithm, and the SINR reciprocal is computed with a Newton-Raphson unit.
What would settle it
Run the fabricated chip in a channel emulator set to the 32×32 WINNER II typical urban micro scenario of Fig. 6(b) and compare measured packet error rates with the fixed-point ASIC curve; if the measured PER at the claimed operating SNR exceeds the simulated curve by more than the design margin, the near-optimal-on-silicon claim is refuted.
Extended reading notes
Core claim
The paper's central claim is that a modified LAMA algorithm, implemented in 28nm CMOS, reaches near-MAP packet error rates for massive multiuser MIMO when the number of users equals the number of base-station antennas, a regime where linear MMSE loses more than 11 dB and channel-hardening message passing fails. The fabricated chip detects 32 users transmitting 256-QAM on the same time-frequency resource, outputs soft information for iterative decoding, and sustains 354 Mb/s at 151 mW in 0.37 mm². The authors state that it is the first silicon prototype of a 32-user massive MU-MIMO detector providing near-optimal error rates under realistic propagation conditions and for symmetric systems. The error-rate curves labeled ASIC are fixed-point simulation results, and the paper claims these match the implemented design closely.
Load-bearing premise
The claim that the chip itself runs close to optimal detection rests on the unverified assumption that the fixed-point simulations labeled ASIC reproduce the hardware exactly, and that the modified algorithm still converges near optimally with only 32 antennas and correlated channels.
Editorial extensions
If this is right
- A base station with only 32 antennas can serve 32 single-antenna users at 256-QAM on the same time-frequency resource while staying close to MAP error rates in realistic channels.
- In the 256×32 regime, LAMA matches linear MMSE error rates without computing a matrix inversion, avoiding the hardware cost of an inverse.
- In the 32×32 typical urban micro channel, the detector gains more than 11 dB over linear MMSE, which can be traded for coverage or higher spectral efficiency.
- The soft-input soft-output interface lets the detector run iterative detection and decoding with an outer channel decoder.
- Voltage-frequency scaling spans 2.66 Mb/s at 123 pJ/bit to 511 Mb/s at higher supply, so one core covers low-power and high-throughput operating points.
Reading between the lines
- The paper does not prove that damped, Gram-based LAMA retains the original algorithm's large-system optimality; the 32×32 and correlated-channel results are empirical, so extending the proof to finite damped systems is an open question.
- The bit-domain max-log mean and variance unit is a reusable idea for any soft-input soft-output receiver that works on bit LLRs, and it could lower area for constellations beyond 256-QAM.
- Until packet-error-rate measurements on the fabricated chip are reported, the ASIC curves in Figs. 5 and 6 are fixed-point simulations; a chip-in-loop channel emulator test would turn the silicon claim into a directly verified one.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript describes a 28nm CMOS ASIC that implements a modified version of the large-MIMO approximate message passing (LAMA) algorithm for massive MU-MIMO uplink detection. The design supports 32 single-antenna users with 256-QAM, provides soft-input soft-output capability for iterative detection and decoding, and is claimed to achieve near-MAP error-rate performance under realistic channel conditions and in symmetric systems with as many users as base-station antennas. The authors report a measured clock frequency of 400 MHz, power consumption of 151 mW, core area of 0.37 mm^2, and a throughput of 354 Mb/s. Error-rate results are presented as packet error rate (PER) versus SNR curves comparing the LAMA implementation against MMSE and CHEMP detectors for 256x32 and 32x32 antenna configurations under Rayleigh and urban-micro channel models.
Significance. If the error-rate claims are fully substantiated, this would be a valuable silicon demonstration: a 32-user, 256-QAM, SISO massive MU-MIMO detector that simultaneously offers 354 Mb/s at 151 mW in a compact 0.37 mm^2 footprint, with more than an 11 dB gain over MMSE in a realistic 32x32 channel. The measured clock, power, and area results are concrete and credible, and the hardware techniques (bit-domain max-log mean/variance computation, Cannon's algorithm for fan-out reduction, pipeline interleaving of two detection problems) are sensible and likely useful to the community. However, the central 'near-MAP' claim is not supported by the evidence presented: the PER curves compare only against MMSE and CHEMP, no MAP or optimality baseline is shown, the only symmetric 32x32 result uses QPSK rather than 256-QAM, and the 'ASIC' PER curves are not explicitly identified as measured or simulated. These gaps directly affect the paper's headline contribution and require correction before the manuscript can be considered for publication.
major comments (4)
- [Section IV, Figs. 5 and 6] The curves labeled 'ASIC' are not explicitly identified as measured packet error rates from the fabricated chip or as results from a bit-true fixed-point simulation. Section IV-A describes measured clock frequency, power, and voltage-frequency scaling, but no test setup for measuring PER on the chip is described. If the PER curves are fixed-point simulation results, the manuscript must state this clearly and relabel the curves (e.g., 'LAMA fixed-point model'), because the current phrasing 'PER of our LAMA ASIC' implies hardware measurements. This distinction is load-bearing for the central claim that the silicon itself provides near-MAP error rates.
- [Sections I and IV, Figs. 5 and 6] The 'near-MAP' label is not supported by the plotted comparisons. The PER curves benchmark only against MMSE and CHEMP; no MAP, max-log MAP, sphere-decoder, or genie-aided optimality bound is provided. An SNR gain of more than 11 dB over MMSE does not quantify the gap to MAP optimality. The theoretical optimality result in [3] is for the large-system limit with i.i.d. Rayleigh fading, and for symmetric systems the introduction itself restricts the optimality claim to QPSK. The authors should either add a finite-dimensional MAP or near-MAP reference for the simulated scenarios or soften the near-MAP wording to 'superior to linear MMSE and CHEMP'.
- [Section IV, Fig. 6] The only symmetric B=U=32 PER results use QPSK with R=0.75; there is no 32x32 256-QAM PER curve anywhere in the manuscript. Since the title and abstract claim a 32-user 256-QAM near-MAP detector for symmetric systems, the manuscript must provide a 32x32 256-QAM fixed-point simulation against a MAP baseline, or explicitly restrict the near-MAP symmetric-system claim to lower-order constellations. Without this, the claimed 256-QAM symmetric near-MAP result is neither demonstrated nor derived.
- [Section II-B and Section IV] The modified LAMA algorithm's key parameters are not reported. The damping factor theta in the message-damping step and the fixed-point word lengths and LUT contents used in the ASIC are not given. The claim that the algorithm achieves near-floating-point performance and near-optimal error rates in realistic channels depends on these parameters, and their absence prevents reproduction and sensitivity analysis. The authors should report theta and the fixed-point configuration, or state clearly that these details are available in an external source.
minor comments (5)
- [Section III-A] The sentence 'Our simulations in Section IV for various antenna configurations and channel models show that our approach entails a negligible performance loss at around 4x lower area' is vague; the specific antenna configurations and the quantified performance loss should be cited here.
- [Algorithm 1] The comment '(message damping details are excluded)' should be clarified: the damped variables tau^t_d are used in line 6 but are not defined in the displayed algorithm, and the relationship between theta, tau^t, and tau^{t-1}_d is only described in the text.
- [Figure 5] The number of LAMA iterations differs between panel (a) (LAMA-8) and panel (b) (LAMA-9); the caption should note that this is intentional and explain the choice.
- [Table I] The normalization formula for 'Normalized Energy' and 'Normalized Area Eff' is not defined in the text; the meaning of the factor (U/32)^2 should be stated explicitly.
- [Abstract] The phrase 'near-optimal error-rate performance' should be defined quantitatively or replaced with a more specific statement about the exhibited gains, because no optimality baseline is shown in the paper.
Circularity Check
No significant circularity; silicon measurements and external comparisons are independent of inputs.
full rationale
The paper's central measured claims—throughput, power, area, and error-rate comparisons—are self-contained hardware results. The LAMA algorithm is cited from the authors' prior theoretical work [3], but that citation is a parameter-free optimality theorem with assumptions (large-system limit, i.i.d. Rayleigh fading) that do not include the target finite-dimensional correlated-channel ASIC implementation. The paper explicitly acknowledges this gap and introduces damping and fixed-point modifications, then validates them in simulation against external baselines (MMSE and CHEMP). The max-log approximation and reciprocal unit are taken from [8], an earlier ASIC paper sharing an author, but they are implementation techniques rather than fitted predictions or renamed outputs. No parameter is fitted to a subset of data and then presented as a prediction; no uniqueness theorem from the same authors is invoked to forbid alternatives; and the 'near-MAP' label, while arguably under-supported by the absence of a MAP baseline, is an evidentiary weakness rather than a circular reduction to the paper's own inputs. The absence of any exhibited equation-level equivalence or fitted-input-as-prediction step means the derivation chain is not circular.
Assumptions & free parameters
free parameters (2)
- Message damping factor theta =
not specified, theta in (0,1]
- Fixed-point word lengths and LUT contents =
not specified
assumptions (4)
- standard math LAMA is MAP-optimal in the large-antenna limit with i.i.d. Rayleigh fading channels, from [3].
- domain assumption The base station has perfect knowledge of the channel matrix H, noise variance N0, and symbol variance Es.
- domain assumption The WINNER II typical urban micro channel model is representative of realistic propagation conditions.
- domain assumption The max-log approximation and Gray-mapping symmetries produce negligible performance loss.
Cite this review
Pith. "Pith review of A 354Mb/s 0.37mm^2 151mW 32-User 256-QAM Near-MAP Soft-Input Soft-Output Massive MU-MIMO Data Detector in 28nm CMOS." pith.science (2026). https://pith.science/paper/5ZCDVLND
@misc{pith2026190803288,
author = {Pith},
title = {Pith review of: A 354Mb/s 0.37mm^2 151mW 32-User 256-QAM Near-MAP Soft-Input Soft-Output Massive MU-MIMO Data Detector in 28nm CMOS},
year = {2026},
howpublished = {\url{https://pith.science/paper/5ZCDVLND}},
note = {Machine review of arXiv:1908.03288}
}
read the original abstract
This paper presents a novel data detector ASIC for massive multiuser multiple-input multiple-output (MU-MIMO) wireless systems. The ASIC implements a modified version of the large-MIMO approximate message passing algorithm (LAMA), which achieves near-optimal error-rate performance (i) under realistic channel conditions and (ii) for systems with as many users as base-station (BS) antennas. The hardware architecture supports 32 users transmitting 256-QAM simultaneously and in the same time-frequency resource, and provides soft-input soft-output capabilities for iterative detection and decoding. The fabricated 28nm CMOS ASIC occupies 0.37mm^2 , achieves a throughput of 354Mb/s, consumes 151mW, and improves the SNR by more than 11dB compared to existing data detectors in systems with 32 BS antennas and 32 users for realistic channels. In addition, the ASIC achieves 4x higher throughput per area compared to a recently proposed message-passing detector.
Figures
Reference graph
Works this paper leans on
-
[7]
A 1.8Gb/s 70.6pJ/b 128× 16 link-adaptive near-optimal massive MIMO detector in 28nm UTBB-FDSOI,
W. Tang, H. Prabhu, L. Liu, V . Öwall, and Z. Zhang, “A 1.8Gb/s 70.6pJ/b 128× 16 link-adaptive near-optimal massive MIMO detector in 28nm UTBB-FDSOI,” in IEEE Int. Solid-State Circuits Conf. (ISSCC) , Feb. 2018, pp. 224–226
work page 2018
-
[3]
Optimality of large MIMO detection via approximate message passing,
C. Jeon, R. Ghods, A. Maleki, and C. Studer, “Optimality of large MIMO detection via approximate message passing,” in Proc. IEEE Int. Symp. Inf. Theory (ISIT) , Jun. 2015, pp. 1227–1231
work page 2015
-
[1]
Massive MIMO for next generation wireless systems,
E. Larsson, O. Edfors, F. Tufvesson, and T. Marzetta, “Massive MIMO for next generation wireless systems,” IEEE Commun. Mag. , vol. 52, no. 2, pp. 186–195, Feb. 2014. softout LLR out LLR in mean/var function unrelated designs onsg. SINR comp. matrix-vector multiplier 32x32 complex 828um 452um Fig. 7. Chip micrograph; LAMA data detector ASIC is highlighted...
work page 2014
-
[2]
An Overview of Massive MIMO: Benefits and Challenges,
L. Lu, G. Y . Li, A. L. Swindlehurst, A. Ashikhmin, and R. Zhang, “An Overview of Massive MIMO: Benefits and Challenges,” IEEE J. Sel. Topics Signal Process. , vol. 8, no. 5, pp. 742–758, Oct. 2014
work page 2014
-
[4]
Channel hardening-exploiting message passing (CHEMP) receiver in large-scale MIMO systems,
T. Narasimhan and A. Chockalingam, “Channel hardening-exploiting message passing (CHEMP) receiver in large-scale MIMO systems,” IEEE J. Sel. Topics Signal Process. , vol. 8, no. 5, pp. 847–860, Oct. 2014
work page 2014
-
[5]
A 0.58mm 2 2.76Gb/s 79.8pJ/b 256- QAM massive MIMO message-passing detector,
W. Tang, C. Chen, and Z. Zhang, “A 0.58mm 2 2.76Gb/s 79.8pJ/b 256- QAM massive MIMO message-passing detector,” in IEEE Symp. VLSI Circuits, Jun. 2016, pp. 1–2
work page 2016
-
[6]
Y . Chen, C. Cheng, T. Tsai, W. Sun, Y . Ueng, and C. Yang, “A 501mW 7.6lGb/s integrated message-passing detector and decoder for polar-coded massive MIMO systems,” in IEEE Symp. VLSI Circuits , Jun. 2017, pp. C330–C331
work page 2017
-
[8]
C. Studer, S. Fateh, and D. Seethaler, “ASIC implementation of soft- input soft-output MIMO detection using MMSE parallel interference cancellation,” IEEE J. Solid-State Circuits , vol. 46, no. 7, pp. 1754–1765, Jul. 2011
work page 2011
Show all 11 references
-
[9]
On the convergence of approximate message passing with arbitrary matrices,
S. Rangan, P. Schniter, and A. Fletcher, “On the convergence of approximate message passing with arbitrary matrices,” in Proc. IEEE Int. Symp. Inf. Theory (ISIT) , June 2014, pp. 236–240
2014
-
[10]
A cellular computer to implement the Kalman filter algorithm,
L. Cannon, “A cellular computer to implement the Kalman filter algorithm,” Ph.D. dissertation, Montana State University, USA, 1969
1969
-
[11]
WINNER II channel models. D1.1.2 V1.2,
P. Kyösti, J. Meinilä, L. Hentilä et al. , “WINNER II channel models. D1.1.2 V1.2,” Tech. Rep. IST-4-027756 WINNER II, 2007
2007
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.