{"id":"278fdf60-7a1b-4e53-84f1-df518d5d8ce9","arxiv_id":"2607.17838","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A blind deep-recurrent-Q-learning controller that steers an energy beam from ternary slot outcomes achieves up to 68% higher throughput than round-robin/random steering and 75-80% of a full-state oracle in simulated WPCNs.","lead":"This paper trains a deep recurrent Q-network to steer an energy beam in a wireless powered network, using only idle/success/collision slot outcomes as feedback, and reports up to 68% throughput gains over non-learning beam steering. The result suggests battery-free IoT devices can be scheduled without channel estimation, charge reports, or device-state tracking.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The simulation's P_th reference is undefined and inconsistent: Eq. (4)+Table I imply M=5 devices rarely exceed P_th=-12 dB, yet Fig. 9 shows abundant transmissions; the 68%/75-80% claims are not reproducible until this is resolved.","rationale":"The central claim—that a transmitter can learn an effective beam-steering policy purely from ternary slot outcomes—is supported only by simulation, so the simulator must faithfully implement the stated physical model. I checked the model equations and found an internal inconsistency that the reader's hardware-determinism concern does not capture. Under the stated parameters, the expected received power for a beam-aligned M=5 device is below the stated P_th, yet the paper reports large numbers of non-idle slots. This suggests the simulation used a normalized threshold (as drawn in Figs. 4/5) rather than the absolute power threshold in Eq. (8), or that some parameter values are misreported. Either way, the quantitative headline figures are not reproducible from the text. This is more load-bearing than the external hardware concern because it does not depend on assumptions about real devices; it questions whether the reported numbers follow from the paper's own equations. I therefore keep the reader's CONDITIONAL verdict: the qualitative idea may be sound, but the numerical claims require clarification and a reproducible simulation before they can be accepted.","tokens_in":16774,"tokens_out":15102,"duration_ms":161891,"concrete_test":"Re-implement the simulator exactly per Eqs. (2)–(8) with Table I parameters, M=5, and P_th=-12 dB interpreted as an absolute power threshold (0.063 W). Log, for each slot, the maximum P_r,i among devices and the number of devices satisfying (8). If the counts do not reproduce the active-slot/collision counts in Figs. 9–10, or if no device ever satisfies (8) while Fig. 9 shows thousands of active slots, the reported gains do not follow from the stated model. A secondary check: rerun with P_th defined relative to the normalized array-factor peak (as Figs. 4–5 suggest) and report which interpretation reproduces Figs. 6, 9, 10.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Equations (2)–(4) and Table I are internally inconsistent. For a device aligned with a beam, the expected squared projection is E|g_i^H w(φ_s)|^2 = σ_l^2 (K M + 1)/(K+1). With P_T = 1 W, σ_l^2 = 10^-2, M = 5, K = 6 dB, this is about 0.042 W = -13.8 dBW, below the stated P_th = -12 dB (0.063 W). Condition (8) would then only be satisfied by rare NLoS fades, but Fig. 9 reports ~1330 non-idle slots per episode at N=125 with M=5 and hundreds of collisions at N=75. Figs. 4/5 draw the threshold at -12 dB relative to the normalized array-factor peak, not relative to P_T as the text states. The paper never defines the reference for P_th. If the implementation used the normalized threshold, the absolute received-power model in Section III is not what was evaluated; if it used the stated absolute threshold, the M=5 results are physically implausible. Either way, the 68% improvement and 75-80% of oracle figures are not reproducible from the equations as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper considers a wireless powered communication network (WPCN) in which an energy transmitter/access point (ET/AP) steers a multi-antenna energy beam among predefined directions to regulate the charging of battery-free energy harvesting devices that transmit under slotted ALOHA. The authors formulate the joint beam-steering and random-access problem as a POMDP, and propose an Action-specific Deep Recurrent Q-Network (ADRQN) that learns the steering policy from only the ternary slot outcome (idle/success/collision), without channel estimation, charge-level reporting, or device-state tracking. A model-predictive oracle policy with full state knowledge is designed as a benchmark. Numerical results report up to 68% throughput improvement over round-robin and random-selection baselines, and 75-80% of the oracle throughput.","tokens_in":17150,"tokens_out":10337,"duration_ms":106504,"significance":"The proposed control loop—regulating access by steering the energy beam based only on slot outcomes—is an interesting and potentially low-overhead approach for WPCNs. The POMDP formulation, the recurrent architecture, and the comparison with an informed benchmark are appropriate first steps, and the paper addresses a real gap in the literature on MAC protocols for energy-harvesting networks. However, the quantitative claims are not currently reproducible: the power-threshold definition is inconsistent with the stated channel model and parameters, the statistical reliability of the simulation results is not documented, and the 'oracle' is a finite-horizon heuristic rather than a proven upper bound. If these issues are resolved, the work would be a useful contribution to the field.","major_comments":[{"comment":"The reference for P_th is inconsistent. Eq. (4) defines P_{r,i}=P_T |g_i^H w(φ_s)|^2. With Table I (P_T=1 W, σ_l^2=10^-2, K=6 dB, M=5), the expected value for a beam-aligned device is E|g_i^H w|^2 = σ_l^2 (K M + 1)/(K+1) ≈ 0.042 W ≈ -13.8 dBW. This is below the stated P_th = -12 dB (≈0.063 W), so condition (8) would be met only by rare favorable fades. Yet Fig. 9 reports ~1330 non-idle slots at N=125 for M=5 and Fig. 12 locates the throughput peak at P_th=-12 dB. Figs. 4/5 instead draw the -12 dB line on a 'Normalized Power' array-factor plot, which is a different reference. The manuscript never defines the reference used in the simulations. Either Eq. (8)/Table I/Fig. 12 are evaluated with an absolute threshold, in which case the M=5 results are physically implausible, or the simulation used a normalized threshold, in which case the absolute received-power model of Section III is not wh","section":"III-B/C/E, Table I, Figs. 4-5, 12"},{"comment":"The number of independent device placements is not stated ('a few independent experiments'), and no error bars, confidence intervals, or per-placement results are reported. Fig. 8 shows that baseline throughput is strongly placement-dependent (e.g., RS gives 0.4992 in setup 1 but 0.2703 in setup 2), so the variance across placements is material. The central claim of a 68% improvement over round-robin/random at N=50 (Fig. 6) and the scaling curves in Figs. 9-10 need to be accompanied by a specification of the number of configurations and a measure of dispersion; otherwise the comparisons cannot be assessed statistically.","section":"VI-A, Figs. 6, 9, 10"},{"comment":"The oracle is not an upper bound. It selects actions via an exhaustive search over a finite horizon k=5 with heuristic tie-breaking (near-threshold parameter ξ), and no argument is given that this equals the optimal fully-observable policy. Calling it a 'performance ceiling' and interpreting the ADRQN-oracle gap as 'the cost of partial observability' is therefore overstated. A finite-horizon informed heuristic is a useful benchmark, but the paper should either demonstrate insensitivity to k and ξ (or provide a bound on the suboptimality), or soften the 'upper bound' language so that the 75-80% figure is understood as a fraction of a specific benchmark policy.","section":"V, Eq. (32), VI-C"}],"minor_comments":[{"comment":"The state dimension is stated as H(S+4), but each history step comprises an S-dimensional one-hot action and a 3-dimensional one-hot observation; the dimension should be H(S+3) unless an additional field is included. Please correct the formula.","section":"IV-C2, Eq. (25)"},{"comment":"Calling the devices 'entirely passive' is imprecise: condition (8) requires each BEHD to measure instantaneous received power and compare it with P_th, in addition to charge-threshold detection. Clarify what hardware/energy cost this entails, or state explicitly that this is an idealized model.","section":"III-E"},{"comment":"The normalization of the power axis in Figs. 4-5 is not defined, and Fig. 12 states 'P_th (in dB relative to P_T)'. These two conventions are inconsistent with the threshold in Table I. Unify the reference point throughout.","section":"VI-B, Figs. 4-5, Fig. 12"},{"comment":"The tie-break parameter ξ=0.85 appears without explanation; give the intuition for the 'fewest devices with charge exceeding ξ·Q_th' criterion.","section":"V-A"},{"comment":"Releasing the simulation code and seeds would materially improve reproducibility of the numerical claims, especially since no code is currently available.","section":"Overall"}],"recommendation":"major_revision","confidential_remarks":"The P_th inconsistency is the main technical obstacle. If the authors can supply a consistent definition and confirm that the simulations used it, the paper could be acceptable after revision; if the absolute threshold was used, the M=5 results need re-evaluation. I do not see this as a reason for reject, because it may be a documentation error, but the numerical claims cannot be verified as written."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one before you cite it. The core idea is genuinely new: an ET/AP that learns to steer its energy beam in a WPCN using only the ternary slot outcome (idle/success/collision), with no channel or battery feedback. That is a valuable addition to the WPCN MAC literature, and the POMDP formulation plus the oracle benchmark make it a clean simulation study. The ablations (RNN vs LSTM vs feedforward) are informative, and the authors are honest about fairness and about the oracle being a heuristic.\n\nThe soft spot is a load-bearing inconsistency in the power threshold. The model in Section III defines P_th as an absolute received power threshold, and Table I gives P_th = -12 dB with P_T = 1 W. For the 5-antenna case, the expected aligned received power from Eq. (4) is about -13.8 dBW, below that threshold. So condition (8) would almost never be met and the devices would rarely transmit. Yet Fig. 9 shows thousands of non-idle slots. The beam coverage plots (Figs. 4-5) draw the threshold relative to the normalized array-factor peak, not to P_T. So either the simulations used a normalized threshold (in which case Section III is not what was evaluated) or the listed parameters are physically implausible. The paper never defines the reference for P_th. This makes the headline 68% improvement and 75-80% of oracle numbers unreproducible from the equations as written.\n\nThere are also the usual simulation-study concerns: no code, no error bars, and the number of device placements is never stated. P_th is tuned on the same metric it is evaluated on, though the paper does disclose the sweep. None of these are fatal on their own, but combined with the P_th issue they mean the quantitative claims need care.\n\nThis deserves a serious referee - the idea is worth engaging with and the writing is clear - but the authors need to fix the P_th definition, re-run or re-report the simulations, and release code before the numbers can be trusted. I'd bring it to a reading group as a case study in how a single undefined reference can undermine a simulation study, but I wouldn't cite it yet.","headline":"Genuinely new WPCN MAC idea, but the P_th reference is undefined and the stated parameters contradict the reported simulations, so the headline numbers aren't reproducible.","tokens_in":17608,"tokens_out":5290,"would_cite":false,"duration_ms":47400,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that an energy transmitter can learn to maximize throughput in a battery-free network solely from coarse slot outcomes, reaching 75-80% of an all-knowing oracle's performance and up to 68% beyond non-learning baselines.","keywords":["wireless powered communication networks","energy beamforming","slotted ALOHA","deep recurrent Q-network","partial observability","battery-free devices","medium access control","throughput maximization"],"falsifier":"A hardware experiment or simulation in which identical charge states and beam directions do not always produce the same slot outcome would falsify the deterministic-state premise and with it the claim that the ternary outcome stream alone is sufficient; for instance, replacing the deterministic P_th rule with probabilistic channel access should erode the reported throughput gains.","tokens_in":16692,"feed_emoji":"⚡","tokens_out":6883,"duration_ms":69785,"temperature":0.7,"pith_summary":"This paper aims to show that a wireless powered communication network can be scheduled without any feedback from its battery-free devices. The energy transmitter/access point steers a narrow energy beam across predefined directions, and by watching only whether each slot was idle, a success, or a collision, it learns which beam to use next. The authors model the problem as a partially observable Markov decision process and solve it with an action-specific deep recurrent Q-network whose hidden state accumulates the history of beams and outcomes. In simulation the learned policy raises throughput by up to 68% over round-robin and random beam selection and reaches 75-80% of an oracle policy that knows every device's charge and channel. If the approach transfers to hardware, it would make medium access control for energy-harvesting IoT networks nearly overhead-free.","feed_headline":"Beam steering alone lifts wireless-powered throughput up to 68%","feed_subtitle":"No channel estimates or charge reports: just slot outcomes, yet it reaches 75-80% of an all-knowing oracle.","key_machinery":"The central object is the ADRQN (action-specific deep recurrent Q-network), a recurrent Q-network whose input is the previous beam index and the current ternary slot outcome, each encoded as a one-hot vector and embedded, then processed by a multi-layer RNN or LSTM. The hidden state accumulates the interaction history and acts as a learned substitute for the unobservable joint charge state. The mechanism is completed by the power threshold P_th, which admits only devices whose received power exceeds it to contend; because the slot outcome is then a deterministic function of beam and state, the history of outcomes carries enough information for the agent to learn to steer.","core_discovery":"The paper's central claim is that beam direction can serve as a medium-access-control knob: by charging different spatial clusters at different rates, the energy transmitter indirectly controls which battery-free devices become eligible to transmit and when. Because a device transmits exactly when its stored charge reaches a threshold and the received power from the current beam exceeds P_th, the slot-level ternary observation is a deterministic function of the hidden joint charge state and the chosen beam. The authors therefore cast beam steering as a POMDP and use an ADRQN—an action-specific deep recurrent Q-network—to compress the history of beam choices and idle/success/collision outcome","pith_inferences":["Editorial inference: if the deterministic transmission rule survives on real hardware, the same outcome-only feedback would likely generalize to multiple energy transmitters and heterogeneous capacitor sizes, since those changes would show up as different charging rhythms in the same ternary history.","Editorial inference: P_th is tuned manually in this paper; a natural extension would be to let the agent co-learn the threshold, turning the idle/collision trade-off into part of the policy rather than an external choice.","Editorial inference: the paper itself acknowledges that fairness is not explicitly guaranteed, which suggests that outcome-only beam steering may need an auxiliary signal or a modified objective before it can serve applications with age-of-information or per-device throughput constraints."],"forward_implications":["Battery-free devices can stay fully passive: no channel estimation, charge reporting, or state tracking is needed, so harvested energy is spent on data rather than protocol overhead.","Beam direction becomes a scheduling knob: by choosing which spatial region charges fastest, the energy transmitter shapes when devices contend without any explicit reservation.","The idle/success/collision history is a sufficient signal for the transmitter to infer hidden charging dynamics, as shown by recurrent architectures outperforming fixed-window DQN and by throughput increasing with history length.","The power threshold P_th lets the network trade idle slots for collisions, with throughput peaking at an intermediate setting (P_th = -12 dB in the simulations).","The gap to the oracle (roughly 20-25% throughput) is the measurable price of partial observability under this protocol."],"fun_headline_variants":["Beam steering alone: 68% throughput boost in WPCNs","No sensing, no reports: beam steering lifts throughput 68%","From idle/success/collision: beam steering gains 68%","Beam steering as MAC: 68% throughput gain"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole learning signal rests on the assumption that a battery-free device transmits exactly when its capacitor is full and the beam's received power exceeds a threshold, so each slot outcome is a clean readout of charging state; probabilistic or imperfect device behavior would break this.","fun_headline_variants_meta":{"raw":{"variants":["Beam steering alone: 68% throughput boost in WPCNs","No sensing, no reports: beam steering lifts throughput 68%","From idle/success/collision: beam steering gains 68%","Beam steering as MAC: 68% throughput gain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000806,"raw_usage":{"total_tokens":3371,"prompt_tokens":734,"completion_tokens":2637,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":2562}},"tokens_in":478,"tokens_out":2637,"duration_ms":21972,"temperature":1.0,"reasoning_tokens":2562,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T16:50:16.698772+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A hardware experiment or simulation in which identical charge states and beam directions do not always produce the same slot outcome would falsify the deterministic-state premise and with it the claim that the ternary outcome stream alone is sufficient; for instance, replacing the deterministic P_th rule with probabilistic channel access should erode the reported throughput gains.","supporting_citations":[],"review_version":1}