REVIEW 3 major objections 5 minor 17 references
Extremely Large Beyond-Diagonal RIS: Low-Rank Modal Optimization for Near-Field Communications
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Near-field geometry makes fully connected beyond-diagonal RIS affordable at extremely large scale by reducing the design to a low-dimensional modal problem.
desk verdict Correct core equivalence and a genuinely useful dimensionality reduction, but the headline entry count quietly depends on a numerical rank tolerance that should be stated up front. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the active subspace of the aperture, $S=\operatorname{span}([G,F^*])$ with $\rho=\dim S\le N+K$; the geometric modal basis $\Phi=[\Phi_S,\Phi_\perp]$ formed by the leading $\rho$ left singular vectors of the stacked atom matrix plus an arbitrary orthonormal complement; the lossless completion $\Theta=\Phi\Psi\Phi^H+(I_M-\Phi\Phi^H)$, which makes every $\Psi\in U(L)$ a valid passive unitary surface; and the Halmos unitary dilation, which supplies the $2\rho\times 2\rho$ block matrix reproducing a given contraction $T=\Pi_S\Theta\Pi_S$ on $S$. Together these turn the $M$-dimensional unitary surface design into an equivalent $L$-dimensional design with $L=2\rho$, and the paper optimizes the latter by unitary steepest descent on $U(L)$ with a factored low-rank line search.
What would settle it
Deploy a $24\times24$ panel at 28 GHz in a scattering-rich indoor environment, or simulate the same geometry with ray-traced multipath, measure the numerical rank $\rho$ of the stacked spherical-wave atom matrix derived from the actual channel, and compare the sum rate optimized at $L=2\rho$ with the fully connected optimum: any gap beyond numerical tolerance, or any $\rho$ significantly larger than $N+K$, would refute the exactness claim.
Extended reading notes
Core claim
The central discovery, stated as Proposition 2, is an exact equivalence: if the active subspace $S=\operatorname{span}([G,F^*])$ of the sampled aperture channel lies inside the span of the chosen modal basis and $L\ge 2\rho$ with $\rho=\dim S$, then the maximum sum rate over $L\times L$ unitary modal matrices equals the maximum over the full $M\times M$ unitary surface. Because in free-space line of sight the sampled base-station-surface and surface-user matrices are exactly collections of spherical-wave atoms at the $N+K$ terminal positions, $\rho\le N+K$, and a basis built from one thin SVD of the stacked atom matrix plus arbitrary orthonormal slack gives $L=2\rho$. The proof uses the Halmos unitary dilation to realize any contraction on $S$ by a $2\rho\times 2\rho$ unitary, embedded in the surface through a lossless completion that leaves the complementary modes reflected unchanged. Consequently the fully connected optimum is attained with $L^2$ tunable entries independent of $M$; in the reference deployment with $N=8$, $K=3$, $M=576$, the measured active dimension is $\rho=7$, so $L=14$ and 196 entries reach the anchor, while the mismatched DFT beamspace needs about 12,100 entries.
Load-bearing premise
The load-bearing premise is that both legs of the cascade are deterministic line-of-sight point-source responses with no random path gains, so the aperture fields are exactly spanned by the known terminal-position wave responses—if multipath, scattering, or position uncertainty adds extra aperture degrees of freedom, the active subspace grows beyond $N+K$ and the promised equivalence at twice that dimension no longer follows.
Editorial extensions
If this is right
- Fully connected BD-RIS optimality becomes reachable at panel-size-independent cost: after a one-time $O(M\rho^2)$ compression, each optimization iteration costs $O(L^3)$ with $L=2\rho$.
- The reconfigurable entry count tracks the propagation geometry, not the aperture: growing the panel from $16\times16$ to $24\times24$ moves the modal operating point only from 144 to 196 entries.
- Far-field DFT beamspace is the wrong coordinate system in the near field: a spherical wavefront atom spreads over roughly $(d_F/2r)^2$ beams, producing about a sixty-fold entry-count penalty in the reference deployment.
- Element-domain group-connected surfaces cannot match the modal design at the same entry budget: they reach 99% of the fully connected anchor only at 18,432 entries in the reference setting, and their per-iteration cost grows with block size.
- The value of beyond-diagonal coupling itself grows with near-field depth: the gap between diagonal phasing and the fully connected optimum widens from 2.5 to 4.4 bit as the base station moves deeper into the near field.
Reading between the lines
- The graceful degradation below $L=2\rho$ visible in the numerical study suggests an adaptive protocol the paper does not state: start from a geometric estimate of $\rho$, optimize at $L=2\rho$, and grow $L$ only when measured rate gains exceed a threshold.
- If position uncertainty or diffuse multipath is introduced, replacing the empirical atom stack by an ensemble covariance turns the geometric basis into a Karhunen-Loève basis, and the exactness condition would become an effective-rank condition rather than the hard $N+K$ count used here.
- In the far-field limit the active dimension collapses toward $K+1$ and a plane-wave beamspace already spans the subspace, so the practical advantage of the geometric modal design is specific to near-field depth and should vanish as the geometry recedes.
- A testable wideband extension is to measure $\rho$ over the joint space-time aperture; one would expect the required $L$ to grow with bandwidth as the spherical-wave atoms become frequency-dependent.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces XL-BD-RIS, a beyond-diagonal RIS architecture for the radiative near field. It models both cascade legs with a free-space Green function, observes that the aperture fields lie in the span of N+K position-dependent spherical-wave atoms, and constructs a compact unitary modal matrix on a subspace of dimension L=2ρ, where ρ is the (numerical) rank of the stacked atom matrix. Proposition 2 claims that this modal representation provably attains the fully-connected unitary-surface optimum with L=2ρ≤2(N+K), independently of the panel size, via a Halmos unitary dilation argument. The paper also develops a WMMSE-Riemannian alternating algorithm with monotone sum-rate convergence and per-iteration cost O(L^3) independent of M, and presents numerical results for a 24×24 panel showing that the fully-connected sum rate is reached with 196 reconfigurable entries, compared with about 12,100 for a DFT beamspace and 331,776 for the fully-connected element-domain surface.
Significance. If the exactness claim is made fully rigorous, the result is significant: it identifies a low-dimensional subspace structure of near-field cascaded channels and converts a prohibitive O(M^2)-entry, O(M^3)-per-iteration problem into an O(L^2)-entry, O(L^3)-per-iteration problem whose dimension depends on the deployment geometry rather than on panel size. The proofs of Propositions 1 and 2 are mathematically sound under the stated exact-rank, LoS, and unitary assumptions, and the WMMSE-Riemannian optimization with its monotonicity result is a solid, standard contribution. The numerical study is extensive and internally consistent, and the comparison against DFT beamspace and group-connected architectures is informative. The main caveats are that the exactness theorem is applied to an ε-truncated channel without an explicit perturbation analysis, and that the claimed equivalence is to the non-reciprocal unitary surface rather than the reciprocal symmetric BD-RIS of the usual literature; these caveats affect the quantitative headline but not the core qualitative idea.
major comments (3)
- [Sec. IV-B, Prop. 2, footnote 2] The central exactness claim is not literally what is proved. Proposition 2 assumes S = span([G,F*]) with ρ = dim S and S ⊆ range(Φ), but the basis construction in Eqs. (32)-(33) sets ρ to the numerical rank of Σ at an unspecified relative tolerance. For the reference deployment (M=576, N=8, K=3), the exact rank of the stacked atom matrix is generically min(M,N+K)=11, so the reported ρ=7 corresponds to a truncated subspace, not to the S used in the theorem. Consequently, the Halmos dilation argument proves equality (35) only for the ε-truncated channel, and the abstract's statement that the design 'provably attains the fully connected optimum with about two hundred entries' overstates what Prop. 2 establishes. The discarded singular values are asserted to be negligible rather than shown to be irrelevant at the operating SNR. Please restate Prop. 2 and the abstract with an explicit ε-truncation caveat, and either use the exact rank (which would give L=22 and 484 entries in this deployment) or provide a quantitative bound or numerical certification that the rate gap caused by the discarded singular values is negligible over the SNR range considered.
- [Sec. II-A, Sec. III-C, footnote 1] The fully-connected optimum in Proposition 2 is the maximum over the full unitary group U(M), whereas lossless reciprocal BD-RIS, which is the standard physical model cited in the introduction, is constrained to symmetric unitary scattering matrices. The paper explicitly retains the non-reciprocal unitary model and leaves symmetry as an extension, but the abstract and the numerical comparisons still refer to 'fully connected BD-RIS' without this qualification. The equivalence to the fully connected optimum should be scoped to non-reciprocal unitary surfaces, and the paper should discuss how the reciprocal-symmetric constraint changes the result. If the authors intend to claim the reciprocal case, the Halmos dilation construction must be adapted to produce symmetric unitary dilations, which is not demonstrated.
- [Sec. IV-A and IV-B] The dimensioning rule L=2ρ depends on two unquantified quantities: the relative tolerance used to define rank_epsilon and the geometry constant η_dof in the heuristic formula rG ≈ η_dof D_BS D/(λ d_BS). The paper states that ρ is 'measured rather than assumed,' but the measurement is threshold-dependent, and the threshold is never specified. The claimed M-independence of the reconfigurable budget is therefore only as meaningful as the chosen tolerance. Please specify the tolerance selection rule, report its numerical value in the experiments, and analyze its sensitivity with respect to the attained rates and the exactness guarantee.
minor comments (5)
- [Sec. VII-A] The reference deployment reports 'the measured rank of G is rG=4' but does not state the relative tolerance used for the numerical rank. Please provide the tolerance and the singular-value spectrum of [G,F*] so that the rank measurement is reproducible.
- [Sec. VII-D, Fig. 5(a)] The paraxial estimate rG ≈ D_BS D/(λ d_BS)+1 uses quantities D_BS and the constant η_dof that are not defined in the main text or the figure caption. Please define all symbols and state the parameter values used for the dashed curve.
- [Sec. VII] The operating point 'SNR is 0 dB' is only loosely defined. Please specify the reference noise power, the normalization of the cascaded channel, and how the per-user noise variance is set in the simulations.
- [Figs. 2 and 3] The label 'Modal (oracle)' is potentially confusing; the modal basis is constructed from known terminal positions, not from an oracle channel. Please clarify in the captions what 'oracle' refers to.
- [General] There are several typographical artifacts in the text, such as 'coefficients' and 'suffices.' A careful proofreading pass is recommended before a revised submission.
Circularity Check
No significant circularity: the M-free modal equivalence is a genuine Halmos-dilation theorem, not a renamed fit or a fitted prediction.
full rationale
The derivation chain is self-contained. In Sec. II-A the LoS model defines the channel matrices as exactly collections of spherical-wave atoms at the terminal positions (Eq. 30); the active subspace S in Eq. 31 is the span of that same atom stack. The geometric basis of Sec. IV-B is the left-singular-vector basis of [G,F*], so the premise S subset of range(Phi) in Prop. 2 holds by construction. This definitional step is not the paper's substantive claim; the substantive step is Prop. 2's use of Halmos dilation ([16], an external result) to show that any contraction on S is realized by a 2rho x 2rho unitary, combined with the lossless completion (26) that preserves passivity. The statement that aperture fields live in span([G,F*]) is a direct consequence of the field definitions, but the equality of the L=2rho optimized rate with the full M-dimensional optimum is an additional theorem, not a restatement of the input. The only self-citation ([14]) is a related-work pointer and is not load-bearing. The numerical-rank thresholding in Sec. IV-B and footnote 2 is a real caveat: exactness is proven on the epsilon-truncated active subspace, while the 196-entry numerical claim uses the measured numerical rank (rG=4, rho=7) rather than the exact rank min(M,N+K)=11. This is a tolerance/correctness concern, not circularity, because the theorem is stated conditionally and the disputed step is not an input renamed as an output. No circular step meeting the quote-and-reduction evidence standard was found; the paper's central claim has independent mathematical content.
Assumptions & free parameters
free parameters (2)
- Numerical rank tolerance epsilon =
unspecified
- Geometry constant eta_dof =
order one, unspecified
assumptions (5)
- domain assumption Propagation is the free-space scalar Green function G(x,y) = exp(-j kappa0 ||x-y||)/||x-y|| with no multipath or random path gains.
- domain assumption Both cascade legs are line-of-sight and the direct BS-user links are blocked.
- domain assumption A passive lossless surface corresponds to a unitary scattering kernel, and any unitary is physically realizable through non-reciprocal elements such as circulators.
- domain assumption The aperture is planar, electrically large, and sampled on a lambda/2 grid so the continuous modal analysis transfers to the discrete element domain.
- standard math Halmos unitary dilation: every contraction on a Hilbert space has a unitary dilation on a space of at most twice the dimension.
Cite this review
Pith. "Pith review of Extremely Large Beyond-Diagonal RIS: Low-Rank Modal Optimization for Near-Field Communications." pith.science (2026). https://pith.science/paper/3A5DBXEW
@misc{pith2026260811476,
author = {Pith},
title = {Pith review of: Extremely Large Beyond-Diagonal RIS: Low-Rank Modal Optimization for Near-Field Communications},
year = {2026},
howpublished = {\url{https://pith.science/paper/3A5DBXEW}},
note = {Machine review of arXiv:2608.11476}
}
abstract
Beyond-diagonal reconfigurable intelligent surfaces (BD-RIS) achieve their best performance when fully connected, at the price of an optimization and hardware burden that grows quadratically, and per iteration cubically, with the number of elements. Extremely large surfaces make this burden prohibitive, while their sheer aperture places both the base station and the users in the radiative near field, where far-field design tools break down. This paper introduces the extremely large BD-RIS (XL-BD-RIS) concept and shows that near-field geometry is precisely what makes fully connected performance affordable at scale. Modeling the cascade with the free-space Green function, we prove that the aperture fields live in a low-dimensional subspace spanned by the spherical-wave responses of the terminal positions, and we design a compact unitary modal matrix on this subspace, built from localization information alone, that provably attains the fully connected optimum with a number of reconfigurable entries set by the geometry and independent of the panel size. A weighted-MMSE Riemannian algorithm optimizes the beamformers and the modal matrix with monotone convergence at panel-size-independent cost. Numerical results show that a $24\times24$-element panel reaches the fully connected optimum with about two hundred entries instead of three hundred thousand. A mismatched DFT beamspace pays a sixty-fold entry penalty rooted in the beam spread of spherical wavefronts, while the classical block-wise architecture delivers strictly lower rates at any matched entry budget.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Modeling and architecture design of reconfigurable intelligent surfaces using scattering parameter network analysis,
S. Shen, B. Clerckx, and R. Murch, “Modeling and architecture design of reconfigurable intelligent surfaces using scattering parameter network analysis,” IEEE Trans. Wireless Commun., vol. 21, no. 2, pp. 1229–1243, 2022
2022
-
[2]
H. Li, S. Shen, and B. Clerckx, “Beyond Diagonal Reconfig- urable Intelligent Surfaces: From Transmitting and Reflecting Modes to Single-, Group-, and Fully-Connected Architectures,” IEEE Trans. Wireless Commun., vol. 22, no. 4, pp. 2311–2324, 2022
work page 2022
-
[3]
——, “Beyond diagonal reconfigurable intelligent surfaces: A multi-sector mode enabling highly directional full-space wireless coverage,” IEEE J. Sel. Areas Commun., vol. 41, no. 8, pp. 2446– 2460, 2023
work page 2023
-
[4]
Closed-form global opti- mization of beyond diagonal reconfigurable intelligent surfaces,
M. Nerini, S. Shen, and B. Clerckx, “Closed-form global opti- mization of beyond diagonal reconfigurable intelligent surfaces,” IEEE Trans. Wireless Commun., vol. 23, no. 2, pp. 1037–1051, 2024
work page 2024
-
[5]
M. Nerini, S. Shen, H. Li, and B. Clerckx, “Beyond diago- nal reconfigurable intelligent surfaces utilizing graph theory: Modeling, architecture design, and optimization,” IEEE Trans. Wireless Commun., vol. 23, no. 8, pp. 9972–9985, 2024
work page 2024
-
[6]
H. Li, M. Nerini, S. Shen, and B. Clerckx, “A Tutorial on Beyond-Diagonal Reconfigurable Intelligent Surfaces: Modeling, Architectures, System Design and Optimization, and Applications,” 2025. [Online]. A vailable: https://arxiv.org/abs/2505.16504
arXiv 2025
-
[7]
W. U. Khan, M. Ahmed, C. K. Sheemar, M. Di Renzo, E. Lagunas, A. Mahmood, S. T. Shah, O. A. Dobre, J. Querol, and S. Chatzinotas, “Survey on Beyond Diagonal RIS Enabled 6G Wireless Networks: Fundamentals, Recent Advances, and Challenges,” arXiv preprint arXiv:2503.08423, 2025
work page Pith review arXiv 2025
-
[8]
Communicating with large intelligent surfaces: Fundamental limits and models,
D. Dardari, “Communicating with large intelligent surfaces: Fundamental limits and models,” IEEE J. Sel. Areas Commun., vol. 38, no. 11, pp. 2526–2537, 2020
work page 2020
Show all 17 references
-
[9]
Channel estimation for extremely large- scale MIMO: Far-field or near-field?
M. Cui and L. Dai, “Channel estimation for extremely large- scale MIMO: Far-field or near-field?” IEEE Trans. Commun., vol. 70, no. 4, pp. 2663–2677, 2022
2022
-
[10]
Beam focusing for near-field multiuser MIMO communications,
H. Zhang, N. Shlezinger, F. Guidi, D. Dardari, M. F. Imani, and Y. C. Eldar, “Beam focusing for near-field multiuser MIMO communications,” IEEE Trans. Wireless Commun., vol. 21, no. 9, pp. 7476–7490, 2022
2022
-
[11]
Communicating with waves between volumes: Evaluating orthogonal spatial channels and limits on coupling strengths,
D. A. B. Miller, “Communicating with waves between volumes: Evaluating orthogonal spatial channels and limits on coupling strengths,” Appl. Opt., vol. 39, no. 11, pp. 1681–1699, 2000
2000
-
[12]
Spatially- stationary model for holographic MIMO small-scale fading,
A. Pizzo, T. L. Marzetta, and L. Sanguinetti, “Spatially- stationary model for holographic MIMO small-scale fading,” IEEE J. Sel. Areas Commun., vol. 38, no. 9, pp. 1964–1979, 2020
1964
-
[13]
Fourier plane- wave series expansion for holographic MIMO communications,
A. Pizzo, L. Sanguinetti, and T. L. Marzetta, “Fourier plane- wave series expansion for holographic MIMO communications,” IEEE Trans. Wireless Commun., vol. 21, no. 9, pp. 6890–6905, 2022
2022
-
[14]
Multi-user holo- graphic communications via channel operator diagonalization,
C. Iacovelli, G. Iacovelli, and S. Chatzinotas, “Multi-user holo- graphic communications via channel operator diagonalization,” IEEE Wireless Communications Letters, vol. 14, no. 9, pp. 2753– 2757, 2025
2025
-
[15]
Capacity maximization for MIMO channels assisted by beyond-diagonal RIS,
E. Björnson and Ö. T. Demir, “Capacity maximization for MIMO channels assisted by beyond-diagonal RIS,” in Proc. 19th Eur. Conf. Antennas Propag. (EuCAP), 2025, pp. 1–5
2025
-
[16]
Normal dilations and extensions of operators,
P. R. Halmos, “Normal dilations and extensions of operators,” Summa Brasiliensis Mathematicae, vol. 2, pp. 125–134, 1950
1950
-
[17]
Steepest descent algorithms for optimization under unitary matrix constraint,
T. E. Abrudan, J. Eriksson, and V. Koivunen, “Steepest descent algorithms for optimization under unitary matrix constraint,” IEEE Transactions on Signal Processing, vol. 56, no. 3, pp. 1134–1147, 2008
2008
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.