{"id":"f62b74f4-c501-4a6a-b08c-103a6f5ac38c","arxiv_id":"2411.11070","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A fuzzy-logic-enhanced multi-agent reinforcement learning framework for joint access point selection, precoding, and reconfigurable intelligent surface phase design is shown in simulation to improve energy efficiency in RIS-aided cell-free massive MIMO.","lead":"This paper proposes a two-layer multi-agent reinforcement learning scheme that jointly chooses which access points serve each user, designs the precoding, and tunes a reconfigurable intelligent surface to maximize energy efficiency in a cell-free massive MIMO network.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"EE magnitudes in Figs. 6–10 are inconsistent with Eq. (13) and the stated parameters: with B=180 kHz, SE≤~22 bps/Hz, and P_CPU=5 W, EE is capped below ~6e5 bit/J, yet plots reach ~1.8e6 bit/J, implying an unstated ~3× bandwidth or power-model change.","rationale":"The reader's weakest assumption pointed to the same issue, and it is indeed the most load-bearing. Other problems—the missing binary mapping for the AP-selection constraint (14d), absent hyperparameters, no code/seeds/error bars, and the fuzzy-logic/federated-learning conflation—are real but secondary; they affect implementation detail or reproducibility, whereas the EE-vs-model discrepancy invalidates the actual numbers in every figure and in the abstract. I checked the arithmetic twice: with SE≈22 bps/Hz (Fig. 9 at M=10), B=180 kHz yields B·SE≈3.96e6 bit/s, and the minimum plausible P_total from Table III is about 6.9 W, so EE≤5.8e5 bit/J. Figures 6, 8, 10 display up to ~1.8e6 bit/J. The gap is a factor of ~3, far beyond rounding or minor parameter ambiguity. A bandwidth of 1.8 MHz or a CPU power of 0.5 W would be needed, contradicting the text. This is not merely 'outside current consensus'; it is an internal inconsistency between the stated model and the presented results. It undercuts the strongest_claim because the 85% improvement is anchored to an unverified absolute scale, and without code or a corrected model one cannot tell whether the relative gains are preserved. I therefore agree with the rejection. The proposed test—recomputing Fig. 8 from Eq. (13) with stated parameters, or obtaining the code to do so—would settle whether the concern lands. If the figures survive the check after a corrected parameter set, the paper might be resubmitted with revised numbers; as is, the central claim is unsupported.","tokens_in":20635,"tokens_out":6270,"duration_ms":59637,"concrete_test":"Ask the authors to provide the simulation script or raw data for Fig. 8 and rerun it with exactly B=180 kHz and P_CPU=5 W; compute the maximum EE via Eq. (13) from the reported SE. If the maximum exceeds about 5.8e5 bit/J, the stated model cannot generate the plot, and the 85% claim needs re-benchmarking.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim is that the proposed FL-based MARL scheme raises EE by 10–85% over baselines. Every plotted point is computed from Eq. (13), EE = B·SE/P_total, with parameters in Section V-A and Table III: B=180 kHz, P_CPU=5 W, P_static_l=50 mW, P_k=10 mW, P_0,l=100 mW, P_element=10 mW, L=8, K=6, N=64. This fixes a tight upper bound. The largest SE in Figs. 7/9/11 is about 22 bps/Hz, giving B·SE ≈ 4.0×10^6 bit/s. Even before adding transmit power, P_total is at least P_CPU + L·P_static + K·P_k + L·P_0 + N·P_element ≈ 5 + 0.4 + 0.06 + 0.8 + 0.64 = 6.9 W (RIS/fronthaul traffic terms are negligible at this bit rate). Hence Eq. (13) allows EE ≤ 4.0×10^6/6.9 ≈ 5.8×10^5 bit/J. Yet Figs. 6, 8, and 10 show EE values up to about 1.8×10^6 bit/J, about 3.1 times higher. Matching the plots would require either B≈1.8 MHz instead of 180 kHz, or P_CPU≈0.5 W and comparable cuts to all static powers, or a different unit axis; the paper states none of these. Because the abstract's 85% figure is taken from the M=10 point of Fig. 8, and the 10%/16% claims come from the same inconsistent curves, the numerical results cannot be validated against the written model. The relative curve ordering might survive a uniform rescaling, but the paper does not show this, and the headline EE improvement is not reproducible from the supplied model.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript considers a downlink RIS-aided cell-free massive MIMO system and formulates a joint precoding, RIS phase-shift, and binary AP-selection optimization problem to maximize energy efficiency. It proposes a double-layer multi-agent deep deterministic policy gradient (MADDPG) architecture, an adaptive power-threshold AP-selection algorithm, and a fuzzy-logic (FL) strategy to reduce the number of agents and accelerate convergence. A complexity comparison is provided, and simulations report that the proposed scheme improves EE by about 10% over full AP coverage at the peak, by 16% and 85% over full-coverage FL-MADDPG and ZF at M=10, and converges in about 45% of the time of plain MADDPG.","tokens_in":20937,"tokens_out":5577,"duration_ms":51777,"significance":"If the numerical results are reliable, the paper would offer a scalable MARL-based solution to a relevant problem in RIS-aided cell-free massive MIMO and a useful complexity reduction. The double-layer decomposition and the explicit complexity comparison in Table I are clear strengths. However, the central quantitative claims are currently undermined by an internal inconsistency between the stated power model/bandwidth and the plotted EE magnitudes, and by missing training details that make the experiments irreproducible. The contribution is therefore not yet established.","major_comments":[{"comment":"The plotted EE magnitudes are inconsistent with the stated parameters and with Eq. (13). With B=180 kHz, a maximum observed sum SE of about 22 bps/Hz (Figs. 7/9/11), and the static powers in Table III (P_CPU=5 W, L*P_static=0.4 W, K*P_k=0.06 W, L*P_0=0.8 W, N*P_element=0.64 W), Eq. (13) gives P_total at least 6.9 W and EE at most B*SE/P_total ~ 5.8e5 bit/J. The EE curves in Figs. 6, 8, and 10 reach about 1.8e6 bit/J, roughly 3.1 times higher. Matching the plots would require B ~ 1.8 MHz, P_CPU ~ 0.5 W, or a different power/unit model, none of which is stated. Because the reported 10%, 16%, and 85% improvements are read from these curves, the numerical claims are not validated by the written model.","section":"V-A, Table III, Eq. (13), Figs. 6-10"},{"comment":"The MARL training details needed for reproducibility are missing: learning rates, batch size, number of training episodes/steps, number of fuzzy agents N_F, membership width d_a, and initialization/update rate tau_threshold. The results are reported without seeds or confidence intervals. Since the paper's claims rest entirely on simulations, these omissions prevent independent verification; please provide a complete hyperparameter table and, if possible, code or data.","section":"IV-B, Algorithm 2, Table II"},{"comment":"The abstract states a general '85% enhancement over the zero-forcing (ZF) method,' but the text reports this figure only at M=10 in Fig. 8, while the peak improvement over full coverage in Fig. 6 is 10% and the conclusion repeats only the 10% figure. The headline claim should be restricted to the configuration actually simulated and should be accompanied by the variability across the settings shown in Figs. 6-12.","section":"Abstract, V-D, Fig. 8"},{"comment":"The mapping from the continuous MADDPG action to the binary AP-selection matrix satisfying (14d) is under-specified. Algorithm 1 uses a power threshold, but Eqs. (15)-(18) treat A as a given matrix, and no discretization step or feasibility-repair mechanism is described. Please specify how binary integer feasibility is enforced during training and execution.","section":"IV-B and constraint (14d)"}],"minor_comments":[{"comment":"Constraint (14b) is written with 'for all k in K' even though it bounds a sum over k; the universal quantifier should be over l (one per AP). In addition, constraint (14c) is called a 'sparse constraint' but it is a phase-shift constraint.","section":"III-C, Eq. (14b)"},{"comment":"The text says 'the concept of Federated Learning (FL) was proposed in [46]', but reference [46] is a book on fuzzy logic and the algorithm subsequently uses FL to mean fuzzy logic. Please correct the terminology and the citation.","section":"IV-A"},{"comment":"The maximum AP transmit power is stated as '5 dB' in the text and '3.2 W' in Table III; please use consistent units (e.g., 5 dBW). Similarly, the figure captions give Pelement = -20 dB, which should be specified as -20 dBW or as 10 mW to match Table III.","section":"V-A and Table III"},{"comment":"The AP-level action space is stated as AAP in L; since the action is an L x K AP-selection matrix, this notation should be corrected to avoid confusion.","section":"IV-B"},{"comment":"The convergence figures have unlabeled axes except for the caption text; please specify what is plotted (e.g., normalized reward versus training episode) and include multiple-seed curves or error bars.","section":"Figs. 3-5"}],"recommendation":"major_revision","confidential_remarks":"The EE magnitude inconsistency is severe: if the authors cannot supply corrected results or code, I would recommend rejection in a subsequent round. The Section IV-A confusion between federated learning and fuzzy logic, and the mismatched citation, should also be checked carefully."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: the paper's headline numbers are not credible as reported, because the simulated EE values contradict the paper's own EE formula and parameter table by about a factor of three. I did the arithmetic: B=180 kHz, SE up to ~22 bps/Hz, P_total at least 6.9 W (just CPU + static + RIS elements), so EE is capped at about 5.7e5 bit/J. Figures 6, 8, 10 show values up to ~1.8e6 bit/J. Either the bandwidth or the power model in the text is wrong, or the figures were generated with a different setup. This is load-bearing: the 10% and 16% improvements and the abstract's \"85% over ZF\" all come from these numbers.\n\nThe genuinely new piece is the architecture: double-layer MADDPG with fuzzy-logic agent aggregation, plus an adaptive power-threshold AP selection, for joint precoding, AP selection, and RIS phases in a user-centric cell-free massive MIMO system. The relative trends look sensible—serving each user from a subset of APs improves EE, and fuzzy aggregation cuts convergence time by about half at a few percent performance cost. The complexity table is a useful contribution, showing that fuzzy aggregation reduces the scaling from O(L^2 ...) to roughly O(L N_F ...). I believe the qualitative story could survive once the energy model is fixed.\n\nSoft spots beyond the magnitude problem: (i) there is no code, no seeds, no error bars, and the hyperparameters are incomplete—Table II gives only two hidden layers, gamma, replay size, and soft update; no learning rates, batch size, episode count, number of fuzzy agents F, membership width d_a, or threshold update rate. (ii) Nowhere is it explained how the continuous actor output becomes the binary AP-selection matrix in constraint (14d); Algorithm 1 describes a threshold but not the mapping. (iii) The abstract's 85% is the M=10 point in Fig. 8, unqualified. (iv) The fuzzy-logic paragraph cites [46], which is actually a fuzzy-logic textbook, and muddles \"Fuzzy Logic\" with \"Federated Learning.\"\n\nWho is it for? Researchers working on RL-based resource allocation in cell-free and RIS networks. They'll find the architecture and complexity analysis interesting, but they should not trust the simulated gains until corrected.\n\nRecommendation: desk reject would be defensible, but I'd send it for review with a strong request for major revision. The combination is new enough, and the flaw looks like a fixable parameter/unit error rather than a fundamentally broken approach. Ask for a corrected simulation model, full hyperparameters, and ideally code.","headline":"Promising algorithmic idea, but the simulated EE values are about three times higher than the paper's own model allows, so the headline gains are not trustworthy as reported.","tokens_in":21677,"tokens_out":6461,"would_cite":false,"duration_ms":56082,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims a fuzzy-logic multi-agent reinforcement learning scheme that selects a subset of access points per user improves energy efficiency over full coverage and zero forcing, with faster convergence.","keywords":["reconfigurable intelligent surface","cell-free massive MIMO","energy efficiency","access point selection","multi-agent reinforcement learning","fuzzy logic","precoding","user-centric networks"],"falsifier":"Reproduce the simulation with B = 180 kHz and the stated static powers; with the plotted sum SE near 14 bps/Hz, Eq. (13) caps EE at about 5×$10^{5}$ bit/J. If the code produces the plotted ~1.8×$10^{6}$ bit/J values, the published system model differs from the simulation, and the gains must be re-evaluated under the stated parameters.","tokens_in":20218,"feed_emoji":"⚡","tokens_out":5999,"duration_ms":59481,"temperature":0.7,"pith_summary":"The paper tries to establish that, in a reconfigurable-intelligent-surface (RIS) aided cell-free massive MIMO network, letting each user be served by a selected subset of access points rather than by all of them is both practical and energy-efficient when the joint choice of precoding, AP selection, and RIS phase shifts is made by a double-layer multi-agent reinforcement learning scheme. The scheme uses an adaptive power threshold to choose which APs serve which users and a fuzzy-logic layer to compress the agents' state space and speed up training. The authors report that, in their simulated configuration, this raises energy efficiency by about 10% at the peak over full coverage and up to 85% over zero-forcing precoding when each AP has 10 antennas, while converging in about 45% of the training time of plain MARL. They also show that increasing transmit power or RIS element count improves spectral efficiency but eventually degrades EE, so the design exposes a trade-off.","feed_headline":"Fuzzy MARL speeds training and beats ZF on energy efficiency","feed_subtitle":"In a RIS-aided cell-free MIMO simulation, AP selection raises peak EE by 10% and up to 85% over ZF.","key_machinery":"The carrying object is the double-layer fuzzy-logic-based MADDPG architecture. In the first layer, each AP is an agent whose action space is the joint precoding vectors and a per-UE binary selection; the selection is produced by an adaptive power threshold that compares each precoding vector's power to a threshold updated toward a target average. In the second layer, the same agents choose RIS phase-shift actions. Fuzzy logic sits between the environment and the agents: membership functions cluster observed states into a smaller number of fuzzy states, and defuzzification maps fuzzy actions back to actual actions, so the network trains on F fuzzy agents instead of L physical APs. This compression is what carries the claimed reduction in convergence time and complexity.","core_discovery":"The central claim is that the EE maximization problem, which is a non-concave mixed-integer program, can be approximated by decoupling it into two coordinated learning layers: a first MADDPG layer that jointly designs precoding and, through an adaptive power threshold, the AP-UE selection matrix; and a second layer that designs the RIS phase shifts given that selection. Fuzzy logic reduces the number of effective agents by mapping the original agent states to a smaller set of fuzzy states through membership functions, cutting computational complexity from exponential to linear in the number of APs at a reported cost of 2.5–3.6% in EE performance. The paper argues that the gain over zero forcing and over full-coverage MADDPG, as well as the faster convergence, support user-centric AP selection and fuzzy acceleration as effective tools for green RIS-aided cell-free networks.","pith_inferences":["The paper's quantitative EE results are hard to reconcile with its stated 180 kHz bandwidth and power model: plugging the reported SE and static power into Eq. (13) gives an EE ceiling of roughly 5×10^5 bit/J, while the plots reach about 1.8×10^6 bit/J, implying the simulator may use a wider bandwidth or a different power model than stated.","The adaptive power threshold acts like a learned sparsification regularizer on the precoding matrix; a similar threshold mechanism could transfer to other massive MIMO resource allocation tasks such as pilot assignment or user scheduling.","The fuzzy-layer idea generalizes: any multi-agent wireless optimization with large state/action spaces could compress agents through membership functions, at a known cost in precision.","The trade-off between RIS element count and element power suggests a design rule: the optimal N decreases as Pelement increases, which could guide hardware choices without full simulation."],"forward_implications":["If the reported gains hold, serving each user by a small learned subset of APs outperforms full-coverage precoding in energy efficiency, especially when transmit power is high.","The fuzzy-logic state compression makes the MARL approach practical at larger AP counts, since complexity scales linearly with the number of APs instead of exponentially.","EE peaks at a moderate transmit power and a moderate number of RIS elements; beyond those points the additional power consumption outweighs the SE gains.","The 2.5–3.6% EE loss of FL-MARL relative to plain MARL is a small price for a 45% reduction in convergence time, so fuzzy compression is a viable deployment trade-off."],"supporting_citations":[{"why":"Supplies the power consumption model that the paper extends to the user-centric RIS-aided setting.","marker":"[42]"},{"why":"Provides the channel model and path loss model adopted in the simulation.","marker":"[21]"},{"why":"Basis for the double-layer MARL architecture with per-AP agents.","marker":"[33]"},{"why":"Source of the fuzzy logic acceleration strategy combined with MARL.","marker":"[35]"},{"why":"Prior deep reinforcement learning approach for AP selection that the proposed method builds on and compares against.","marker":"[24]"},{"why":"Shows MARL is used for AP clustering in cell-free systems, motivating the MARL-based AP selection here.","marker":"[34]"}],"fun_headline_variants":["Fuzzy MARL boosts RIS-aided cell-free MIMO EE by 85%","Joint precoding and AP selection: fuzzy MARL beats ZF by 85%","User-centric AP selection and fuzzy MARL: 85% energy gain","Speedy fuzzy MARL for green RIS-aided cell-free networks","Fuzzy acceleration in MARL cuts complexity, lifts EE 85% over ZF"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The EE and SE curves rest on the bandwidth and power-consumption parameters stated in Section V-A and Table III; if the simulator actually used different values, the reported percentage improvements are unanchored.","fun_headline_variants_meta":{"raw":{"variants":["Fuzzy MARL boosts RIS-aided cell-free MIMO EE by 85%","Joint precoding and AP selection: fuzzy MARL beats ZF by 85%","User-centric AP selection and fuzzy MARL: 85% energy gain","Speedy fuzzy MARL for green RIS-aided cell-free networks","Fuzzy acceleration in MARL cuts complexity, lifts EE 85% over ZF"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000211,"raw_usage":{"total_tokens":1456,"prompt_tokens":1025,"completion_tokens":431,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":641,"completion_tokens_details":{"reasoning_tokens":327}},"tokens_in":641,"tokens_out":431,"duration_ms":4932,"temperature":1.0,"reasoning_tokens":327,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:57:43.077724+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reproduce the simulation with B = 180 kHz and the stated static powers; with the plotted sum SE near 14 bps/Hz, Eq. (13) caps EE at about 5×$10^{5}$ bit/J. If the code produces the plotted ~1.8×$10^{6}$ bit/J values, the published system model differs from the simulation, and the gains must be re-evaluated under the stated parameters.","supporting_citations":[{"cited_title":"On the total energy efﬁciency of cell-free massive MIMO,","cited_arxiv_id":null,"evidence_quote":"Supplies the power consumption model that the paper extends to the user-centric RIS-aided setting."},{"cited_title":"A joint precoding framework for wid eband reconﬁgurable intelligent surface-aided cell-free netwo rk,","cited_arxiv_id":null,"evidence_quote":"Provides the channel model and path loss model adopted in the simulation."},{"cited_title":"Double-laye r power control for mobile cell-free XL-MIMO with multi-agent rein forcement learning,","cited_arxiv_id":null,"evidence_quote":"Basis for the double-layer MARL architecture with per-AP agents."},{"cited_title":"Uplink Power Control for Extremely Large-Scale MIMO with Multi-Agent Reinforcement Learning and Fuzzy Logic","cited_arxiv_id":"2302.09290","evidence_quote":"Source of the fuzzy logic acceleration strategy combined with MARL."},{"cited_title":"Energy efﬁcient AP selection for cell-free massive MIMO sy stems: Deep reinforcement learning approach,","cited_arxiv_id":null,"evidence_quote":"Prior deep reinforcement learning approach for AP selection that the proposed method builds on and compares against."},{"cited_title":"Access point clustering in cell-free massive MIMO using conventio nal and federated multi-agent reinforcement learning,","cited_arxiv_id":null,"evidence_quote":"Shows MARL is used for AP clustering in cell-free systems, motivating the MARL-based AP selection here."}],"review_version":1}