{"id":"658179d4-eb67-44e5-9742-90d382906eda","arxiv_id":"2504.13455","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A three-stage compressed-sensing and weighted-least-squares position estimator exploits modular XL-array geometry in Terahertz systems and outperforms simulated baselines with lower complexity.","lead":"This paper designs a three-stage algorithm that estimates 3D user positions in Terahertz systems with modular extra-large antenna arrays. It uses compressed sensing for angle estimation and weighted least squares for position refinement, claiming better simulated accuracy and lower complexity than existing baselines.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Visible-SA selection and LoS identification assume NLoS power is negligible; threshold ψ and reflection coefficient Γ are never varied, so the SNS-aware low-SINR claim is not established.","rationale":"The paper's central claim is conditional on the SNS stage working: Stage 1 selects visible SAs using normalized receive power and identifies LoS as the largest dictionary correlation. The authors explicitly assume NLoS loss is much higher than LoS (Section III-B) and 'much higher channel gain of the LoS path' (Algorithm 1). This is not a tuning detail; it determines which anchors feed the WLS and the Stage 3 dictionary center. The simulation section never reports ψ or sweeps Γ, and Fig. 10 prescribes VR masks, so it tests robustness to a given VR layout rather than the ability of Eq. (17) to discover the VR. Without such a test, the claimed low-SINR advantage over NF-JCEL and DFT-MUSIC could be an artifact of a favorable channel draw. A secondary correctness issue is the apparent z-axis phase-sign inconsistency in Eq. (8) with respect to the steering vector in Eqs. (9)-(10), but the reader's identified power-dominance assumption is the more load-bearing concern for the empirical claim. The proposed sweep would either expose the failure or confirm that the LoS/NLoS separation is robust across realistic parameter ranges.","tokens_in":21615,"tokens_out":11066,"duration_ms":109963,"concrete_test":"Rerun the default simulation while sweeping Γ (e.g., 0.1 to 1.0) and ψ (e.g., 0.1 to 0.9), keeping the scatterer positions from Table I, and report: (i) visible-SA classification accuracy relative to the ground-truth χ masks; (ii) final RMSE at SINR = -10 dB and 0 dB; (iii) whether the coarse estimate used to center the Stage 3 dictionary remains within the ±8-grid-step window. If RMSE or classification accuracy degrades sharply for any realistic Γ/ψ combination, the central robustness claim needs substantial qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the power-based SNS handling: visible-SA selection in Section III-B via Eq. (17) and the one-block SOMP LoS identification in Algorithm 1, Step 4 both rely on LoS paths carrying far more power than NLoS paths and on invisible SAs being separable by a normalized-power threshold ψ. The paper states this as an assumption ('the loss of NLoS paths is much higher than that of the LoS path') but never reports ψ, never sweeps the LoS/NLoS power ratio controlled by the reflection coefficient Γ in Eq. (5), and the VR experiments in Fig. 10 force visibility patterns a priori rather than testing whether Eq. (17) discovers them. If a strong single-reflection NLoS path, or a low-power-but-visible SA, crosses the threshold, the wrong anchors enter the WLS in Stage 2, biasing the coarse position and therefore the Stage 3 reduced-dictionary center. Because the claimed low-SINR advantage in Fig. 6 is precisely the regime where power fluctuations and noise are largest, this assumption is the least secure part of the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper considers 3-D localization of multiple single-antenna UEs in a THz uplink system equipped with a modular XL-MIMO array with sub-connected hybrid beamforming. The authors model the channel with a hybrid spherical-planar wave model, where each sub-array sees a distinct incident angle while a planar wavefront is assumed within each SA, and incorporate spatial non-stationarity via visibility indicators. They propose a three-stage algorithm: (i) identify visible SAs based on normalized received power and estimate AoAs of a few typical visible SAs via simultaneous orthogonal matching pursuit (SOMP) over a frequency-domain block-sparse dictionary; (ii) obtain a coarse UE position via iterative weighted least squares (WLS) from the AoA pseudo-linear equations; (iii) estimate AoAs of the remaining visible SAs with a geometrically reduced dictionary and refine the position with WLS. Simulations compare RMSE versus SINR against a near-field joint channel estimation and localization benchmark and a collocated DFT-MUSIC design, and examine the effects of SA interval, antenna allocation, training blocks, and visibility region layout. The central claim is that the proposed framework achieves accuracy close to the full-dictionary upper bound at substantially reduced complexity, and outperforms the benchmarks especially at low SINR.","tokens_in":21764,"tokens_out":11774,"duration_ms":101989,"significance":"If the claims are upheld, the paper contributes a practical and computationally efficient pipeline for SNS-aware 3-D localization in modular XL-MIMO THz systems. The use of geometric priors to shrink the CS dictionary is a sensible complexity-reduction idea, and the modular-array versus collocated comparison is a useful design insight. The WLS pseudo-linear derivation is correct to first order, and the three-stage coarse-to-fine structure is well motivated. However, the significance is tempered by the fact that the visible-SA selection and LoS-identification steps rest on an untested power-separation assumption, and the z-axis phase convention in Eq. (8) appears to be the conjugate of the standard UPA response; both need to be resolved before the simulation results can support the paper's conclusions.","major_comments":[{"comment":"The visible-SA selection criterion in Eq. (17) assumes that LoS received power at visible SAs 'highly exceeds' that at invisible ones and that a fixed threshold ψ separates them. The paper never states the value of ψ used, never sweeps ψ, and never varies the reflection coefficient Γ in Eq. (5), which controls the NLoS power level. The VR experiments in Section VI-F impose visibility patterns a priori rather than testing whether Eq. (17) discovers them from received powers. Since a strong single-reflection NLoS path or a low-power visible SA could place an invisible SA above the threshold or a visible one below it, the selected anchors in Stage 1 can be wrong, biasing the WLS coarse estimate and the Stage 3 reduced-dictionary center. This directly undermines the claimed low-SINR advantage in Fig. 6. The authors should report the default ψ, provide a sensitivity study over ψ and Γ, and ideally test the selection when the LoS/NLoS power gap is reduced.","section":"Section III-B, Eq. (17) and Fig. 10"},{"comment":"The virtual elevation AoA is defined as φ = -sinϕ in Eq. (8), so the z-axis steering vector in Eq. (10) has phase e^{-j2π d/λ sinϕ}, the conjugate of the standard UPA response for a wave arriving from elevation ϕ. With the array deployed along the positive z-axis and the geometry in Eq. (3b) defining ϕ as the elevation above the x-y plane, a physical plane wave from that direction should exhibit a phase advance e^{+j2π d/λ sinϕ}. If the channel model in Eq. (6) and the dictionary both use the flipped sign, the system is internally consistent but unphysical; if any step uses the standard convention, the estimated ϕ is negated and the WLS z-coordinate estimate is biased. The authors must verify the sign convention against the physical array response, correct Eqs. (8) and (10) if needed, and re-run the simulations with the corrected model.","section":"Section II-B, Eqs. (8) and (10)"}],"minor_comments":[{"comment":"The ToA estimator derived via MUSIC in Eqs. (28)-(32) is not used in either Stage 2 or Stage 3; the positioning pipeline relies solely on AoA. Either integrate the ToA as an additional measurement or remove this section to avoid a dangling contribution.","section":"Section III-D"},{"comment":"The weight matrix W depends on the covariance Rz of the AoA estimation errors, but the paper does not specify how Rz is computed or estimated. Algorithm 2 updates W using this covariance, but no model or empirical procedure is given. Please state the error covariance used in the simulations or provide an approximation.","section":"Section IV-B, Eq. (45)"},{"comment":"The text says 'D = 0.2 m in default simulation setup' but Table I lists the default SA interval as D = 1 m; clarify which value is used for the low-interval case.","section":"Section VI-C"},{"comment":"The reference to 'Fig. 8' in the discussion of training blocks should be 'Fig. 9'.","section":"Section VI-E"},{"comment":"The schemes SOMP-LS and OMP-LS are plotted but never defined in the text; add a brief description of these benchmarks.","section":"Table II and Fig. 6"},{"comment":"There are several typographical errors, including 'structual' in the abstract and 'indicting' in Section VI-C; the paper would benefit from careful proofreading.","section":"Abstract and throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope as a signal-processing contribution, though the positioning focus may be more aligned with communications or vehicular technology venues. The authors should also clarify the novelty relative to the near-field localization works [41] and [42], especially regarding the geometric dictionary reduction, which is the main algorithmic departure. The two major comments above are load-bearing; in particular, the sign convention in Eq. (8) should be double-checked against the simulation code, as the reported RMSE values would be difficult to obtain if the code follows the text literally."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a legitimate simulation paper, not a paradigm shift. The three-stage combination—power-based visible-SA selection, SOMP AoA estimation on typical SAs, WLS coarse positioning, then geometry-reduced dictionary SOMP on the rest—is genuinely new in the near-field localization literature, and the pseudo-linear WLS derivation in (33)–(38) checks out. I verified the residual approximations are first-order correct; the distance terms cancel as they should. The complexity reduction from pruning the dictionary is real and is the paper's most useful contribution. The simulation section supports the central claim conditionally: fine estimation tracks the full-dictionary upper bound and beats the two benchmarks at low SINR.\n\nThe soft spots are where the reader says they are. The visible-SA selection and LoS identification lean entirely on the assumption that LoS power at visible SAs 'highly exceeds' NLoS and invisible-SA power. The threshold ψ in (17) is never reported and never swept, and the reflection coefficient Γ in (5) is fixed, so the LoS/NLoS power ratio is never varied. The Fig. 10 VR experiments force visibility patterns a priori rather than test whether (17) discovers them. That matters because Stage 2 triangulates only from the selected SAs; if a strong NLoS path or a low-power visible SA crosses the threshold, the coarse estimate is biased and Stage 3 inherits the bias. The low-SINR advantage in Fig. 6 is exactly the regime where this assumption is most fragile.\n\nThe z-axis phase convention in (8) with φ = −sinϕ also looks physically inverted relative to the geometry in (3). The paper never discusses it. The system is self-consistent, so the relative comparison likely survives, but the absolute elevation estimates would mirror if the convention is wrong. This needs a sentence or a sign fix.\n\nMinor but worth saying: no Monte Carlo count, no error bars, and the default operating point (D = 1 m, K = 25, VR = 5×5) is a favorable regime. The paper's own Fig. 7 shows performance degrades sharply at D = 0.2 m and D = 2 m, so the 'satisfactory accuracy with evident complexity reduction' claim is tied to that setup. The citation pattern is normal; the HSPWM prior work is appropriately cited, and the self-citations are not inflated.\n\nBottom line: the algorithm is coherent, the math mostly holds up, and the new combination deserves referee time. I would send it to review, but ask for a threshold sweep, MC statistics, a phase-convention check, and at least one NLoS-heavy scenario. It is not a desk reject, and it is not a design recipe yet.","headline":"A coherent three-stage modular-XL-MIMO localization algorithm with a real dictionary-pruning complexity win, but the SNS-handling step rests on an untested power-threshold assumption and the simulations need tightening before the results become a design recipe.","tokens_in":22411,"tokens_out":4106,"would_cite":true,"duration_ms":40375,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A modular extra-large Terahertz array with one RF chain per sub-array can localize users in three dimensions by combining spherical-wave geometry between sub-arrays with planar-wave compressed-sensing angle estimation inside them.","keywords":["3-D localization","Terahertz","modular XL-MIMO","hybrid spherical-planar wave model","spatial non-stationarity","compressed sensing","SOMP","weighted least squares"],"falsifier":"Set up the same simulation with an added strong single-reflection NLoS path whose received power at some sub-arrays is comparable to the LoS power, then compare the sub-arrays selected by the threshold in (17) with the true line-of-sight visibility set. If misclassified sub-arrays displace the WLS estimate or the RMSE rises sharply relative to the LoS-only baseline, the power-separation premise fails; the paper does not sweep the threshold $\\psi$ or the LoS and NLoS power ratio.","tokens_in":21311,"feed_emoji":"📡","tokens_out":9580,"duration_ms":82616,"temperature":0.7,"pith_summary":"The paper is trying to establish that a modular extra-large Terahertz antenna array, in which each of many small sub-arrays has a single radio-frequency chain, can deliver 3-D user localization without the prohibitive complexity of full near-field processing. Its proposed three-stage method selects the sub-arrays that actually see a user, estimates their line-of-sight angles by compressed sensing, triangulates those angles with weighted least squares, and then refines the angles of the remaining visible sub-arrays on a small dictionary built around the coarse position. The central claim, backed by simulations in the paper, is that the fine estimate lands close to the full-dictionary accuracy while cutting complexity, and that it beats a collocated array design and a near-field joint-estimation benchmark at low signal-to-noise ratios. A sympathetic reader would see this as evidence that spatial non-stationarity, usually a nuisance in extra-large arrays, can be turned into an advantage by using only the informative sub-arrays as localization anchors.","feed_headline":"Three-stage method puts 3-D positioning on modular Terahertz arrays","feed_subtitle":"By using only the visible sub-arrays and a reduced dictionary, it keeps near-full accuracy with far less computation.","key_machinery":"The central object is the hybrid spherical-planar wave model (HSPWM), which assigns each sub-array its own pair of azimuth and elevation angles of arrival (spherical-wave relation to the user) while letting all antennas inside a sub-array share one steering vector (planar-wave approximation). It does the load-bearing work: the spherical side makes every visible sub-array a geometrically distinct anchor for WLS triangulation, and the planar side keeps the per-sub-array angle estimation cheap enough for compressed sensing. The three algorithmic mechanisms that run on top of it are the normalized-power visibility test, the SOMP block-sparse recovery of the LoS angle, and the reduced dictionary that shrinks each non-typical sub-array's search space from $I_k J_k$ to $(2\\bar{i}+1)(2\\bar{j}+1)$ codewords around the coarse position.","core_discovery":"Using the hybrid spherical-planar wave model (HSPWM), the paper treats the channel to each sub-array as a planar-wave steering vector whose antennas share one azimuth and one elevation angle of arrival, while the sub-array locations themselves are tied to the user position through spherical-wave geometry. On this model the localization problem becomes: find which sub-arrays have line-of-sight paths, estimate their angles, triangulate, and refine. The paper's contribution is a complete pipeline that does this: a normalized received-power criterion selects visible sub-arrays; simultaneous orthogonal matching pursuit over subcarriers, formulated as block-sparse recovery, estimates the LoS angles of the strongest visible sub-arrays; pseudo-linear equations feed those angles into an iterative weighted least squares coarse position; and a reduced dictionary centered on that coarse position estimates the angles of the remaining visible sub-arrays before a final WLS refinement. Simulation results show the fine RMSE close to the full-dictionary upper bound and better low-SINR performance relative to the benchmarks, with complexity dominated by $O(K_{\\mathrm{Ref}} I_k J_k M_S N I)$ rather than by processing every sub-array with a full dictionary.","pith_inferences":["The paper estimates time of arrival with a MUSIC step but never feeds those ranges into the WLS estimator; combining ToA ranges with the AoA-based pseudo-linear equations is a natural next step that should help when few sub-arrays are visible.","The visible and non-visible split depends on a fixed normalized-power threshold that is never swept; a calibration curve of RMSE versus the threshold under varying LoS and NLoS power ratios would test how much margin the selection step actually needs.","Because the reduced dictionary is centered on the coarse position, a coarse estimate far from the true position could bias the fine stage; an adaptive multi-resolution dictionary that re-expands when the coarse residual is large would guard against that failure mode.","The optimal SA interval and AE allocation results suggest the array layout can be treated as a design variable: one could pose a joint layout-and-estimation optimization that maximizes positioning accuracy under a hardware budget, an optimization the paper stops short of solving."],"forward_implications":["Using only a small number of strongest visible sub-arrays for the first estimate, then a reduced dictionary for the rest, the fine position lands close to the full-dictionary result while cutting per-stage complexity by a quadratic factor.","Because sub-arrays with no line-of-sight path are discarded, the method tolerates spatial non-stationarity: with only nine of twenty-five sub-arrays visible, positioning remains acceptable and approaches the no-blockage case at high SINR.","The sub-array interval has an optimal value: larger spacing widens the angular spread the WLS triangulation sees, but beyond a point the far sub-arrays fade enough that their angle estimates hurt rather than help.","For a fixed total antenna count, there is a best split between the number of sub-arrays and antennas per sub-array, since sub-arrays supply anchors and antennas supply per-anchor angular resolution.","Increasing training blocks improves RMSE but with diminishing returns, suggesting transmit power is a more effective way to buy accuracy than longer pilots."],"supporting_citations":[{"why":"Provides the modular XL-MIMO architecture of sub-arrays with inter-SA spacing that the paper takes as its system model.","marker":"[31]"},{"why":"Justifies the use of spherical-wave relations between sub-arrays and planar waves within each sub-array.","marker":"[32]"},{"why":"Shows the hybrid spherical-planar wave model in use for cross-field channel estimation, the modelling precedent this paper adapts to localization.","marker":"[39]"},{"why":"The near-field XL-MIMO joint activity, channel, and location estimator used as Benchmark 1, whose AoA outputs the paper recombines with its WLS method.","marker":"[41]"},{"why":"Supplies the THz multi-ray channel model, including the LoS and NLoS path-loss and reflection-coefficient expressions used by the simulations.","marker":"[43]"},{"why":"Supplies the molecular absorption coefficient model for the 200-400 GHz band used to set frequency-dependent path loss.","marker":"[45]"}],"fun_headline_variants":["Modular THz arrays get fast 3D positioning via visible sub-arrays","Hybrid wave model cuts computation for Terahertz 3D localization","Sparse recovery and reduced dictionary speed up THz positioning","Three-stage pipeline boosts THz array location accuracy and speed","Visible sub-arrays alone enable fast THz 3D positioning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method hinges on the assumption that a sub-array with a line-of-sight path receives much more power than one without it, so a fixed normalized-power threshold cleanly separates visible from invisible sub-arrays; if a reflected path or shadowing blurs that separation, the wrong anchors are selected and the triangulation is biased.","fun_headline_variants_meta":{"raw":{"variants":["Modular THz arrays get fast 3D positioning via visible sub-arrays","Hybrid wave model cuts computation for Terahertz 3D localization","Sparse recovery and reduced dictionary speed up THz positioning","Three-stage pipeline boosts THz array location accuracy and speed","Visible sub-arrays alone enable fast THz 3D positioning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00064,"raw_usage":{"total_tokens":3007,"prompt_tokens":1063,"completion_tokens":1944,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":679,"completion_tokens_details":{"reasoning_tokens":1853}},"tokens_in":679,"tokens_out":1944,"duration_ms":11848,"temperature":1.0,"reasoning_tokens":1853,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:09:07.933560+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Set up the same simulation with an added strong single-reflection NLoS path whose received power at some sub-arrays is comparable to the LoS power, then compare the sub-arrays selected by the threshold in (17) with the true line-of-sight visibility set. If misclassified sub-arrays displace the WLS estimate or the RMSE rises sharply relative to the LoS-only baseline, the power-separation premise fails; the paper does not sweep the threshold $\\psi$ or the LoS and NLoS power ratio.","supporting_citations":[{"cited_title":"Multi-user modular XL-MIMO communications: Near-field beam focusing pattern and user grouping,","cited_arxiv_id":null,"evidence_quote":"Provides the modular XL-MIMO architecture of sub-arrays with inter-SA spacing that the paper takes as its system model."},{"cited_title":"Near-field modeling and performance analysis of modular extremely large-scale array commu- nications,","cited_arxiv_id":null,"evidence_quote":"Justifies the use of spherical-wave relations between sub-arrays and planar waves within each sub-array."},{"cited_title":"Cross-field channel esti- mation for ultra massive-MIMO THz systems,","cited_arxiv_id":null,"evidence_quote":"Shows the hybrid spherical-planar wave model in use for cross-field channel estimation, the modelling precedent this paper adapts to localization."},{"cited_title":"Sensing user’s activity, channel, and location with near-field extra-large-scale MIMO,","cited_arxiv_id":null,"evidence_quote":"The near-field XL-MIMO joint activity, channel, and location estimator used as Benchmark 1, whose AoA outputs the paper recombines with its WLS method."},{"cited_title":"Multi-ray channel modeling and wideband characterization for wireless communications in the terahertz band,","cited_arxiv_id":null,"evidence_quote":"Supplies the THz multi-ray channel model, including the LoS and NLoS path-loss and reflection-coefficient expressions used by the simulations."},{"cited_title":"A distance and bandwidth dependent adaptive modulation scheme for THz commu- nications,","cited_arxiv_id":null,"evidence_quote":"Supplies the molecular absorption coefficient model for the 200-400 GHz band used to set frequency-dependent path loss."}],"review_version":1}