{"id":"976d6efd-1735-4cb4-aa92-b7b16ccdad1a","arxiv_id":"2507.01727","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A particle-filter-based dual-control framework auto-learns unknown wave parameters online and generates a harmonic power take-off profile that approaches the theoretical maximum energy output for regular waves and shows robustness in one irregular-wave simulation.","lead":"This paper designs a two-level controller for wave energy converters that learns unknown ocean wave parameters online and sets the power take-off force to maximize average energy. The authors test it in MATLAB simulations against model predictive control, extremum seeking, and Bang-Bang control, and report faster adaptation and higher energy capture.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed outperformance of MPC rests on a deliberately short 0.25 s prediction horizon; a properly tuned MPC may reverse the comparison.","rationale":"The paper's analytical core is mostly coherent: the average-power expression (18) follows from the linear model (2) under single-frequency excitation, and the optimality conditions (19)-(22) are consistent with complex-conjugate control. The advertised novelty is the DCEE-based auto-optimization that is claimed to outperform existing controllers, and the most load-bearing support for that claim is the comparison in Table 2. That comparison uses an MPC with a 0.25 s prediction horizon, only 5% of the 5 s wave period, so it does not provide a meaningful test of MPC performance. This is not a disagreement with consensus; it is a correctness risk to the empirical conclusion. I also note a secondary issue: Eq. (34) states that solving (34) and (33) give the same optimal average energy, but this is not generally true when Pavg has posterior variance because E[(Pmax - P)^2] = (Pmax - E[P])^2 + Var(P), so the DCEE cost (40) explicitly trades mean against variance. That is consistent with the dual-control intent, but it means the active-learning controller is not maximizing expected energy at every step and the theoretical framing is looser than stated. The proposed concrete test, running MPC with an adequate horizon, would settle whether the headline comparison is real. If MPC then produces equal or higher energy, the paper must weaken its abstract or use a different benchmark; if DCEE still wins, the concern is resolved. Either way, the reader's CONDITIONAL verdict remains appropriate, so no verdict change is recommended.","tokens_in":15463,"tokens_out":11281,"duration_ms":137267,"concrete_test":"Re-run the steady-state regular-wave comparison in Section 4.2 (Table 2) with the MPC prediction horizon extended to at least one full wave period (5 s; 500 steps at 0.01 s), using the same WEC parameters, constraints (Fu,max, xmax, vmax), and MPC formulation from reference [41]. If the resulting MPC average energy reaches or exceeds DCEE's 13.95 W, the abstract's outperformance claim is unsupported. Also repeat the irregular-wave comparison with this MPC and report at least 10 independent trials with means and standard deviations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 2 in Section 4.2 reports DCEE at 13.95 W versus MPC at 11.65 W. The MPC implementation described in Section 4.2 uses a 0.01 s control interval and a 25-step horizon, i.e. 0.25 s, against a 5 s wave period (omega = 0.4 pi in Section 4.1). A horizon of 5% of the wave period cannot approximate the reactive optimal control that MPC is designed to provide, and the text itself concedes that longer horizons improve energy. The abstract's blanket claim of outperforming model predictive control is therefore not established by this benchmark; the comparison is closer to DCEE versus a truncated one-step-lookahead rule. In addition, all results are single-run with no error bars, so the reported 2.3 W margin cannot be assessed for significance. The regular-wave assumption in Eq. (3) is a modeling limitation that the authors acknowledge, but the decisive empirical gap is the baseline configuration.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-level autonomous control framework for wave energy converters (WECs). A high-level DCEE (Dual Control for Exploration and Exploitation) controller estimates unknown wave parameters (amplitude, phase, frequency) with a particle filter and generates a sinusoidal PTO force profile, parameterized by amplitude, phase, and frequency. The key analytical result is an expression for average generated power as a function of the PTO parameters and the wave parameters, Eq. (18), from which the closed-form optimal PTO profile is derived, Eqs. (19)-(22). A low-level controller is assumed to track this harmonic force profile. Simulations compare the proposed scheme against extremum seeking, MPC, and Bang-Bang control under regular and irregular waves. The claimed contributions are the extension of DCEE to non-stationary periodic optimal conditions and the active-learning-based auto-optimization architecture.","tokens_in":15646,"tokens_out":4107,"duration_ms":49436,"significance":"If the analytical derivation is made rigorous and the simulation benchmark is made fair, the paper would offer a useful design-oriented parametrization of WEC optimal control: it converts a non-stationary control problem into a parameter-learning problem, provides closed-form optimal PTO parameters with no fitted constants, and explicitly quantifies estimation uncertainty for active learning. The extension of DCEE to periodic optimal operating points is conceptually novel. However, the current evidence is weakened by a singularity at the resonance condition used for the optimality derivation, an MPC baseline with an extremely short prediction horizon, and single-run simulations with no statistical assessment. The paper is internally consistent within its linear regular-wave model, but that model is also used by the controller, the estimator, and the simulated plant, so the simulations mainly demonstrate self-consistency rather than independent validation.","major_comments":[{"comment":"Eq. (17) is singular at omega_u = omega because the denominator contains (omega_u - omega), yet the optimality condition (19) selects exactly omega_u = omega. The subsequent derivation of the optimal amplitude (22) substitutes the optimal frequency and phase into the singular expression, which is formally not defined. A limiting argument is needed: for fixed T, as omega_u -> omega, the term (cos(phi) - cos(phi + (omega_u - omega)T)) / (2T(omega_u - omega)) tends to sin(phi)/2, and with the phase condition (21) it tends to 1/2, yielding the positive term Au hex A / (2 sqrt((m omega - K/omega)^2 + h_r^2)). Without this limit, the derivation of (22) and P*_avg in (27) is not rigorous.","section":"Section 2.3, Eqs. (17)-(22)"},{"comment":"The claim of outperforming model predictive control is not established by the presented MPC baseline. The MPC uses a control interval of 0.01 s and a prediction horizon of 25 steps, i.e., 0.25 s, while the regular wave period is 5 s (omega = 0.4 pi in Section 4.1). A horizon of 5% of the wave period cannot approximate the reactive optimal control that MPC is designed to provide, and the text itself acknowledges that longer horizons improve performance. The abstract's blanket statement that the proposed method outperforms MPC should be either removed or substantiated with a properly tuned MPC baseline with a horizon on the order of the wave period or longer.","section":"Section 4.2, Table 2, and abstract"},{"comment":"All simulation results are single runs with no error bars or multiple noise seeds. The reported margin between DCEE (13.95 W) and MPC (11.65 W) is 2.3 W, but without repeated trials or confidence intervals it is impossible to assess whether this margin is statistically significant, especially given the 5% Gaussian measurement noise. The claim of 'effectiveness and robustness' under irregular waves would be much stronger with an ensemble of simulations under different noise realizations and wave seeds.","section":"Section 4, Figures 5-9 and Table 2"},{"comment":"The controller, the particle-filter likelihood, and the simulated plant all use the same linear model and the same average-power expression Eq. (18). The irregular-wave simulation still relies on the dominant-harmonic assumption for the DCEE controller, while the low-level tracking loop is delegated to reference [33] and is not simulated or verified here. Consequently, the simulations demonstrate that the DCEE algorithm is self-consistent with its internal model, but they do not validate that the proposed harmonic PTO force profile maximizes energy in a real broadband sea or that the low-level controller can track the harmonic force with sufficient accuracy. The authors should either include a higher-fidelity plant model that does not share Eq. (18) or explicitly position the results as model-based feasibility rather than physical validation.","section":"Section 3.1, Section 3.5, and Eq. (18)"}],"minor_comments":[{"comment":"The update line 'theta_u,k = theta_u,k-1 + [Delta theta_u; alpha_k Delta theta_u]' is malformed; the notation should be unified with Eq. (35) and the preceding random-step description, e.g., theta_u,k = theta_u,k-1 + alpha_k Delta theta_u.","section":"Algorithm 2"},{"comment":"The sentence 'Upper alpha_min and lower alpha_max bounds' has the words Upper and lower swapped; it should read 'lower alpha_min and upper alpha_max bounds.'","section":"Section 3.4"},{"comment":"The sentence 'The generated Power Take-Off (PTO) profile as the reference for the low-level physical system to follow' is missing a verb; it should be 'serves as the reference.'","section":"Abstract"},{"comment":"The y-axis label 'Energy (W)' is dimensionally inconsistent; if the quantity is average power, the label should be 'Power (W)' or 'Energy (J)' if integrated over 60 s.","section":"Figure 8, lower-right panel"},{"comment":"The row 'Computational consumption' lists '0.34s 42.68s \\'; the backslash appears to be a placeholder and the units should be stated consistently, e.g., seconds per 60-s simulation.","section":"Table 2"},{"comment":"The phrase 'in which the two new average energy durations T1, T2 are set as ( T1 < T2 < T)' is grammatically awkward and should be rephrased for clarity.","section":"Section 3.5"}],"recommendation":"major_revision","confidential_remarks":"The paper is largely self-consistent and the analytical derivation is algebraically coherent once the resonance limit is supplied. The main barriers to acceptance are the unfair MPC baseline, the lack of repeated simulations, and the fact that the validation loop uses the same model for plant, estimator, and controller. These are fixable within the manuscript's scope, so I recommend major revision rather than rejection. The authors should also carefully reword the abstract claim about outperforming MPC."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [colleague],\n\nThis paper's genuine contribution is the transfer of the DCEE active-learning framework to WEC operation with a trigonometric parametrization of the PTO force. The derivation of the average-power expression (18) is clean, the optimality conditions (19)-(22) correctly reduce to known complex-conjugate control for monochromatic waves, and the two-level architecture with particle-filter learning is coherent and well explained. I would not quibble with Section 2.\n\nThe soft spots are empirical and concentrated in Section 4. The Table 2 comparison uses an MPC with a 0.25 s prediction horizon against a 5 s wave period, which is 5% of the wave cycle. That MPC cannot perform reactive optimal control, and the text itself concedes that longer horizons improve energy. So the claimed outperformance over MPC is not established by this benchmark; it is closer to DCEE versus a truncated one-step lookahead. All quantitative results are single-run with no error bars, and no code is provided, so the reported 2.3 W margin cannot be assessed. The irregular-wave test shows full-information MPC producing more energy, which undercuts the abstract's blanket claim. There is also an unaddressed singularity at ωu = ω in Eq. (17), exactly the resonance the optimality condition selects; the limiting value is what delivers the optimal energy. The particle-filter likelihood is never stated, which hurts reproducibility.\n\nThe paper is honest about the regular-wave design assumption, and the citation pattern is appropriate, including the authors' own DCEE papers. The circularity concern—same model in the plant, estimator, and controller—is real but not disqualifying for a first demonstration; it means the simulations show self-consistency rather than independent validation. A serious revision should include Monte Carlo trials, a properly tuned MPC with a realistic horizon and wave preview, and full implementation details.\n\nWho this is for: control researchers in WEC or in DCEE/active learning who want a clean parametric formulation and a clear architecture for non-stationary periodic operation. They will get value from Sections 2 and 3. The empirical claims need work before they are usable. I would send this to peer review rather than desk-reject, but the referees should insist on the baseline fix and repeated trials.","headline":"Useful DCEE-for-WEC paper with a correct parametric derivation, but the MPC benchmark is set to fail and the single-run results do not support the abstract's claims.","tokens_in":16210,"tokens_out":2772,"would_cite":false,"duration_ms":31840,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a wave energy converter in unknown seas can self-optimize by learning just three wave parameters and applying a closed-form PTO profile, and demonstrates that the resulting two-level controller outperforms MPC…","keywords":["wave energy converter","dual control for exploration and exploitation","active learning","particle filter","power take-off optimization","auto-optimization control","non-stationary optimal operation"],"falsifier":"Run the proposed controller in a numerical wave tank with a measured broadband spectrum (for example, a JONSWAP spectrum with a broad peak) and compare the actual average power against the value predicted by Eq. (18) using the estimated dominant frequency, phase and amplitude. If the prediction error is large enough that the DCEE cost function no longer ranks the candidate actions correctly, the central claim fails. A second check is to verify that the low-level loop actually tracks the commanded harmonic PTO force within the stated constraints when the high-level profile changes.","tokens_in":15226,"feed_emoji":"🌊","tokens_out":4258,"duration_ms":50165,"temperature":0.7,"pith_summary":"The paper's central claim is that a wave energy converter operating in unknown, changing seas can be made to self-optimize by recognizing that the optimal PTO force is a single sinusoid whose frequency, phase, and amplitude are determined by the incoming wave's frequency, phase, and amplitude. Once that parametrization is accepted, maximizing average energy reduces to learning three unknown wave parameters and then applying closed-form formulas for the PTO profile. The authors build a two-level controller: a high-level DCEE layer runs a particle filter to estimate the wave parameters and chooses PTO settings that both improve expected energy and reduce estimation uncertainty, while a low-level loop tracks the harmonic force. Simulation comparisons suggest this active-learning scheme converges faster than extremum seeking and yields more energy than MPC and Bang-Bang control under the tested regular and irregular waves.","feed_headline":"Self-learning controller tunes wave energy harvesters to peak power","feed_subtitle":"It learns wave frequency, phase and amplitude on the fly and beats MPC, extremum seeking and Bang-Bang control in simulation.","key_machinery":"The load-bearing object is Eq. (18), the average-power law that connects PTO profile parameters θu = (Au, Bu, ωu) to wave parameters θ = (A, B, ω). Eq. (18) has two terms: a negative mechanical-impedance term from the PTO acting on its own induced velocity, and a cross term from the PTO interacting with the wave-induced velocity. The paper uses this law three ways: to derive the closed-form optimum (19)-(22), as the measurement model in a particle filter that estimates θ from average energy observations, and as the predictor inside the DCEE cost function, whose two terms reward moving toward the estimated optimum and probing to shrink estimation variance. The random step size αk is added to avoid local minima in the non-convex power surface.","core_discovery":"The paper claims that for a point absorber in a regular wave, the average generated power over a horizon T is given by the function Pavg(θu, θ) in Eq. (18), built from the known hydrodynamics and the three PTO parameters Au, Bu, ωu and three wave parameters A, B, ω. Maximizing this function gives the optimal PTO profile in closed form: match the wave frequency (ωu* = ω), set the phase according to Eq. (21), and set the amplitude according to Eq. (22). The maximum average power then depends only on wave amplitude. The paper further claims that by wrapping this formula in a particle filter and a DCEE search over 27 candidate profile adjustments, the controller can learn the unknown wave parameters online, with the DCEE cost function explicitly rewarding actions that reduce future estimation uncertainty, and that the resulting system outperforms the compared benchmarks.","pith_inferences":["A natural extension is to replace the single-sinusoid PTO profile with a truncated Fourier series; the same average-power derivation would produce a set of optimal harmonic coefficients, and the DCEE search would then operate over more parameters.","The single-frequency assumption means the comparison against MPC under irregular waves likely favors the proposed method whenever the irregular sea has a strong dominant harmonic; in genuinely broadband seas the Eq. (18) model may be biased, and a testable extension is to weight Eq. (18) by the wave spectrum.","The active-learning logic could be reused for other periodic energy systems whose optimal operating point moves, such as tidal turbines or oscillating water columns, whenever the reward-per-period is a smooth function of a few phase-amplitude parameters."],"forward_implications":["If the wave is a single-frequency sinusoid, the PTO should oscillate at exactly the wave frequency; any frequency mismatch reduces average energy through the cross-term in Eq. (18).","The maximum average power under the optimal profile is A^2 hex^2 / (8 hr), so once the wave amplitude is learned, the achievable energy ceiling is known without further optimization.","Because the DCEE cost includes an uncertainty term, the controller will intentionally deviate from the current best PTO profile when that deviation yields informative energy measurements.","The high-level design operates on a slow timescale over many wave periods, so the low-level loop must track the harmonic force; the paper delegates that tracking to an existing method and assumes it succeeds."],"supporting_citations":[{"why":"Supplies the DCEE framework that the high-level controller is built on.","marker":"[17]"},{"why":"Supplies the active-learning perspective that motivates minimizing estimation uncertainty.","marker":"[18]"},{"why":"Supplies the hydrodynamic model of the point absorber and the MPC formulation used as a baseline.","marker":"[12]"},{"why":"Provides the extremum seeking control benchmark used in the dynamic learning comparison.","marker":"[6]"},{"why":"Supplies the MPC implementation and simulation baseline for steady-state comparison.","marker":"[41]"},{"why":"Supplies the Bang-Bang control baseline used in the steady-state comparison.","marker":"[42]"},{"why":"Supplies the WEC simulation parameters and hydrodynamic coefficients used in all tests.","marker":"[40]"},{"why":"Supplies the constrained particle filter method that the paper uses for wave parameter estimation.","marker":"[37]"}],"fun_headline_variants":["Active learning boosts wave energy converter output in simulation","Self-tuning wave energy controller adapts to unknown seas","Wave energy harvester learns to match wave frequency for max power","Adaptive control maximizes wave energy capture with minimal probing","Wave energy converter auto-optimizes via active learning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The design assumes the sea is a single-frequency sine wave and the PTO force is a single sine wave of the same shape, so every learned optimum and every predicted measurement follows from Eq. (18); if real waves have energy spread across many frequencies, or the low-level controller cannot track the harmonic force, the learned profile is no longer optimal.","fun_headline_variants_meta":{"raw":{"variants":["Active learning boosts wave energy converter output in simulation","Self-tuning wave energy controller adapts to unknown seas","Wave energy harvester learns to match wave frequency for max power","Adaptive control maximizes wave energy capture with minimal probing","Wave energy converter auto-optimizes via active learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00015,"raw_usage":{"total_tokens":1198,"prompt_tokens":946,"completion_tokens":252,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":175}},"tokens_in":562,"tokens_out":252,"duration_ms":3408,"temperature":1.0,"reasoning_tokens":175,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:45:16.922974+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the proposed controller in a numerical wave tank with a measured broadband spectrum (for example, a JONSWAP spectrum with a broad peak) and compare the actual average power against the value predicted by Eq. (18) using the estimated dominant frequency, phase and amplitude. If the prediction error is large enough that the DCEE cost function no longer ranks the candidate actions correctly, the central claim fails. A second check is to verify that the low-level loop actually tracks the commanded harmonic PTO force within the stated constraints when the high-level profile changes.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the DCEE framework that the high-level controller is built on."},{"cited_title":"Chen, Perspective view of autonomous control in unknown envi- ronment: Dual control for exploitation and exploration vs reinforcement learning, Neurocomputing 497 (2022) 50–63","cited_arxiv_id":null,"evidence_quote":"Supplies the active-learning perspective that motivates minimizing estimation uncertainty."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the hydrodynamic model of the point absorber and the MPC formulation used as a baseline."},{"cited_title":"Parrinello, P","cited_arxiv_id":null,"evidence_quote":"Provides the extremum seeking control benchmark used in the dynamic learning comparison."},{"cited_title":"Khedkar, A","cited_arxiv_id":null,"evidence_quote":"Supplies the MPC implementation and simulation baseline for steady-state comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Bang-Bang control baseline used in the steady-state comparison."},{"cited_title":"Zhang, S","cited_arxiv_id":null,"evidence_quote":"Supplies the WEC simulation parameters and hydrodynamic coefficients used in all tests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the constrained particle filter method that the paper uses for wave parameter estimation."}],"review_version":1}