{"id":"11ab9c81-d7d2-48ac-bb90-390da9ca06ad","arxiv_id":"2412.12400","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"In a simulated spasmodic walleye fishery, machine-learned harvest rules beat conventional rules only for trophy fishing, by waiting for big year classes to grow, and mean fish weight helps only in that setting.","lead":"Fishery managers usually set fishing limits from stock biomass alone, even for fish populations with unpredictable boom years. This study tests machine-learned and conventional harvest rules on a simulated Alberta walleye fishery and finds that adding a mean-weight observation only helps when the goal is trophy fish, not maximum or stable catch.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PPO convergence of the two-observation policies is the load-bearing assumption: the yield/HARA negative results could reflect under-training rather than the uninformative value of mean weight.","rationale":"The reader's weakest_assumption points to PPO convergence, and the manuscript's own text weakens it: Section 3 explicitly says the 2RL HARA rule is noisy and that longer training would likely converge to UMSY. That is a textbook case of a negative result that may be an artifact of optimization failure. The paper is otherwise well structured, the simulation is described in detail, and the companion code enables re-training; the trophy result is likely robust because the effect size is large. A CONDITIONAL verdict is therefore appropriate: the central claim should be accepted only after the proposed convergence diagnostics are reported. I do not see a basis for REJECT, since the authors are transparent about the limitation, and I would not ACCEPT as-is because the headline claim overstates confidence in a negative result that rests on a single RL run per setting. My agreement is with the reader's weakest_assumption; I add the architecture-size confound and the lack of multiple seeds as concrete ways to settle it.","tokens_in":14639,"tokens_out":6442,"duration_ms":58124,"concrete_test":"Retrain the 1RL and 2RL policies for the yield and HARA scenarios with at least 10 random seeds each, the same 6-million-step budget, and evaluation checkpoints every 500k steps; also train a 1RL policy with the [256,64,16] network and a 2RL policy with the [64,32,16] network to separate observation value from architecture/training difficulty. If best-of-seed 2RL yield utility is within ~1 standard error of 1RL and best-of-seed 2RL HARA utility no longer trails oPP/1RL, the negative claim is supported. If either improves materially (e.g., 2RL HARA approaches or exceeds 400), the paper's conclusion that mean weight is uninformative for yield/HARA would need to be weakened to 'no benefit under the original training protocol'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The conclusion that mean weight is useful only for trophy utility depends on the 2RL policies being near-optimal for yield and HARA. Section 2.6 reports one PPO run per policy (6 million steps) with no learning curves, multiple seeds, or best-of-seed selection. Section 3 concedes that the 2RL HARA policy is noisy and that a moving average, larger networks, and longer training could smooth it toward UMSY; that admission is exactly the failure mode that would invalidate the negative result. In Table 2, yield 2RL (249.47 ± 23.86) is not better than 1RL (250.25 ± 22.82), and HARA 2RL (391.27 ± 21.51) is worse than 1RL (402.43 ± 21.14) and oPP (401.73 ± 20.93). Those gaps are consistent with an under-optimized 2RL policy, especially since the 2RL network is larger ([256,64,16] vs [64,32,16]) and may need more data to converge. The trophy advantage (126.90 vs 96.44) is large enough to survive some under-training, so the positive finding is less threatened; the 'surprisingly not useful' claim for yield/HARA is the fragile part. A statement that a policy is 'within one standard deviation' is not a test of equivalence or of optimization convergence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses a simulated age-structured walleye fishery with spasmodic recruitment and partial observability to compare harvest control rules: a constant exploitation rate (U*), a precautionary policy calibrated from the model's optimized U* and B*, an optimized precautionary rule, and neural-network policies trained with PPO using either biomass alone or biomass plus mean weight. Policies are evaluated under yield-maximizing, HARA risk-averse, and trophy-fishing utilities. The central empirical claims are that an additional mean-weight observation improves harvest control rule performance only for the size-dependent trophy utility, where the two-observation RL policy attains roughly 30% more utility than the other policies (Table 2: 2RL 126.90 vs UMSY 96.44), and that mean weight is not useful for yield or HARA utility.","tokens_in":14921,"tokens_out":10253,"duration_ms":88871,"significance":"If the claims hold, the paper makes a useful contribution by showing, in a realistic POMDP-style fishery simulator, that the value of an additional state observation depends strongly on the management objective. It also reproduces a non-intuitive pulse-harvest strategy for trophy objectives, with a clear mechanism shown in Figures 8 and 9. Strengths include a well-specified age-structured simulation model, an MSE-style evaluation protocol, and companion open-source code. The positive trophy result is large and internally consistent. However, the negative result on mean weight for yield and HARA is currently fragile because it rests on the assumption that the PPO-trained 2RL policies are near-optimal; the manuscript itself provides evidence that the HARA 2RL policy is not converged.","major_comments":[{"comment":"The central negative result, that the mean-weight observation is not useful for yield or HARA utility, is not established because the two-observation PPO policies may be under-optimized. Only one PPO run per policy configuration is reported in Section 2.6, with no learning curves or seed variation, and Section 3 concedes that the HARA 2RL policy is noisy and would plausibly converge to UMSY with longer training or smoothing. In Table 2, HARA 2RL (391.27 ± 21.51) is worse than 1RL (402.43 ± 21.14) and oPP (401.73 ± 20.93), and yield 2RL (249.47 ± 23.86) is no better than 1RL (250.25 ± 22.82); these patterns are exactly what one would expect from an under-trained larger network rather than from the uninformative value of mean weight. Please provide convergence diagnostics such as learning curves, multiple seeds, or best-of-seed selection, or a formal near-optimality test, before concluding that mean weight is not useful for these objectives.","section":"Section 2.6 and Section 3, Table 2"},{"comment":"The claim that 'nearly all policies obtain essentially equal amounts of utility' rests on overlapping standard deviations, which is not a statistical equivalence test. The phrase 'within a standard deviation' in the table caption is not sufficient evidence that yield and HARA performances are equal or near-optimal. Report paired differences with confidence intervals, using the shared stochastic sequence design already applied in Figures 5-9, or use an equivalence test with a pre-specified margin.","section":"Section 3, Table 2"},{"comment":"The statement that 'considerable gains can be achieved' by optimizing harvest control rule parameters overstates the results in Table 2: the large gain is confined to the trophy-utility 2RL policy (126.90 vs 88.51-96.44), while for yield and HARA the optimized policies are close to or worse than UMSY (e.g., yield 2RL 249.47 vs UMSY 237.28; HARA 2RL 391.27 vs UMSY 400.58). Please qualify the abstract and introductory framing to distinguish the strong trophy result from the null yield/HARA result.","section":"Abstract and Section 1"}],"minor_comments":[{"comment":"Typo: 'systes' should be 'systems'.","section":"Section 1"},{"comment":"Typo: 'velieve' should be 'believe'.","section":"Section 5"},{"comment":"The productivity variable π is defined but never used subsequently; either use it or remove the definition.","section":"Equation (23)"},{"comment":"The notation 'argmaxF' is unclear; the subscript or function name should be defined explicitly.","section":"Equation (22)"},{"comment":"The 'conventional precautionary policy' (cPP) is not an independent DFO default: its parameters are derived from this paper's optimized U* and B* via Eq. (17). The text should state this explicitly wherever cPP is described as the recommended government of Canada rule.","section":"Section 2.5"},{"comment":"The reference list has inconsistent formatting, including missing parentheses around years and incomplete journal details, and several in-text citations such as 'Sutton, Barto, and others 1998' should follow the journal's style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"I do not see a circularity problem in optimizing and evaluating on the same simulator; that is standard MSE practice. The main risk is optimization convergence: the negative claim requires either demonstrating 2RL convergence or reframing the conclusion as 'with this training budget, mean weight did not help.' The trophy result appears robust and is the strongest contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful applied RL paper for fisheries, but the claim that mean weight only helps for trophy utility is only as good as the PPO training diagnostics, and those are thin.\n\nWhat's new: most prior RL-for-fisheries work uses simple models or single objectives. Here they set up a partially observed age-structured walleye model with spasmodic recruitment, compare three utility functions (yield, HARA risk-averse, trophy), and systematically ask whether adding a mean-weight observation improves the policy. That comparison is new. The methods are solid: the population model is reasonable, the observation error is explicit, and they optimize standard reference-point policies with Bayesian optimization and RL policies with PPO. Code is on GitHub. The trophy result is the real contribution: 2RL gets about 30% more trophy utility than everything else, and the pulse-fishing explanation is convincing and connects to classic theory (Walters 1969). That part I believe.\n\nSoft spots: the negative result for yield and HARA is fragile. Table 2 shows 2RL yield (249.47) is not better than 1RL (250.25), and 2RL HARA (391.27) is actually worse than 1RL (402.43) and oPP (401.73). The authors themselves note the 2RL HARA policy is noisy and that a moving average, larger networks, or longer training might make it converge to UMSY. That is exactly the failure mode that would explain why mean weight appears useless: the 2RL policy may simply be under-optimized. There are no learning curves, no multiple seeds, no best-of-seed selection, and only one PPO run per policy. So the claim 'mean weight is not useful for yield/HARA' is not yet established. At best it's a provisional negative result that needs convergence diagnostics and equivalence testing (or at least paired comparisons with proper uncertainty). The abstract also overstates the gains from optimization; Table 2 shows large gains only in the trophy column, not for yield or HARA.\n\nAlso, the cPP is derived from the same model's U* and B*, so it's not an independent benchmark. That's minor and standard in MSE, but worth noting.\n\nWho it's for: people working on RL for natural resource management, and fisheries scientists interested in whether additional observations are worth collecting. The paper is a good case study and a cautionary example of how RL can discover non-intuitive policies (pulse fishing) but also how easy it is to over-interpret negative results without proper training diagnostics.\n\nRecommendation: deserves peer review, but should be revised to either add PPO training diagnostics (multiple seeds, learning curves) or soften the negative claims. If the authors can show the 2RL policies are converged, the paper becomes much stronger. If not, the honest conclusion is 'we didn't find evidence that mean weight helps for yield/HARA in this model,' not 'mean weight is not useful.'","headline":"Useful RL-for-fisheries case study, but the headline negative result about mean-weight observations is only as strong as the PPO convergence diagnostics, which are currently thin.","tokens_in":15485,"tokens_out":2643,"would_cite":true,"duration_ms":23193,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Mean fish weight helps harvest rules only when the goal is trophy-sized catch, not yield or stable harvests.","keywords":["harvest control rules","reinforcement learning","spasmodic recruitment","walleye fishery","partial observability","mean weight observation","trophy fishing utility","age-structured population model"],"falsifier":"Re-run the HARA scenario with the same 2RL architecture but with an order of magnitude more training steps and a moving-average mean-weight observation; if that policy then exceeds the constant-U policy's utility, the claim that mean weight has no value for risk-averse management would be overturned.","tokens_in":14381,"feed_emoji":"🎣","tokens_out":10396,"duration_ms":88549,"temperature":0.7,"pith_summary":"This paper asks whether extra information—mean fish weight on top of stock biomass—improves harvest control rules in a simulated age-structured walleye fishery with spasmodic recruitment and observation error. The authors find that the value of that extra observation depends entirely on the management objective. For yield maximization and for a risk-averse utility that penalizes variable catches, a rule using only biomass performs as well as any rule using mean weight. For a trophy-fishing objective that values only large fish, the mean-weight observation lets a reinforcement-learned rule wait for strong cohorts to reach trophy size and then harvest in pulses, gaining about 30% more utility than constant-exploitation or precautionary rules. If correct, this means monitoring priorities and policy flexibility should be matched to what managers actually value.","feed_headline":"Mean-weight data helps only when trophy fish are the goal","feed_subtitle":"A simulated walleye fishery shows the extra observation pays off for trophy-sized catch, not for yield or stable harvests.","key_machinery":"The central object is a partially observed Markov decision process (POMDP): a 20-age-class Beverton-Holt walleye model whose true state is the biomass per age class, but an agent sees only noisy observations of survey-vulnerable biomass $B^{survey}$ and mean fish weight $\\bar{W}^{survey}$, both with multiplicative Gaussian error. Recruitment is spasmodic by construction, with a 2.5% probability per year of a pulse 10–30 times average. The mechanism that carries the argument is the neural-network harvest control rule trained by proximal policy optimization (PPO), which can express nonlinear, multi-observation feedback rules without a pre-specified functional form. In the trophy scenario the network learns a bang-bang pulse policy: harvest rate near zero while the dominant cohort is small, then a spike of $U \\approx 0.75$ when mean weight and biomass are high. Mean weight is the informative signal because a recruitment pulse produces a characteristic dip in mean weight followed by a rise as the cohort ages into survey vulnerability and trophy size.","core_discovery":"In a 20-age-class stochastic model of an Alberta walleye fishery calibrated to measured spasmodic recruitment (a 2.5% yearly chance of a pulse 10–30 times average recruitment), the paper compares five harvest control rules under three utilities: total yield, HARA risk-averse utility ($U^{\\gamma}$ with $\\gamma=0.6$), and trophy utility that counts only fish older than age 10. Policies were a constant exploitation rate $U_{MSY}$, a conventional precautionary rule (cPP), an optimized precautionary rule (oPP), a neural-network rule using only survey biomass (1RL), and a neural-network rule using biomass plus mean weight (2RL). The central finding is that the 2RL rule's extra observation does not help under yield or HARA utility—all optimized policies perform nearly identically, with only the conventional precautionary rule clearly worse under HARA—but it delivers a clear gain under trophy utility (mean 126.90 vs 96.44 for $U_{MSY}$ and 92.73 for 1RL). The learned trophy policy is a pulse-harvesting strategy: it avoids fishing while a large cohort is young and mean weight is low, then applies a short high-exploitation pulse around years 8–15 after the recruitment event, when that cohort has reached trophy size. The paper interprets this as evidence that age-structure information becomes decision-relevant precisely when the utility function is size-dependent, and that flexible policy search is most valuable in that regime.","pith_inferences":["A testable extension: the value of the mean-weight observation should increase as the trophy size threshold rises, so re-running the comparison with thresholds between ages 5 and 15 would map where the advantage appears.","The paper's own suggestion that longer training would push the 2RL HARA policy toward $U_{MSY}$ implies that the negative result could be an optimization artifact; a much longer training run is the direct experiment to settle that.","A regulator-friendly approximation to the learned trophy policy would be a hand-coded pulse rule: skip or sharply cut fishing for a fixed window after a detected large recruitment event, then reopen at high exploitation; if such a rule captures most of the 2RL gain, complex neural policies may not be needed in practice.","If pulsed trophy harvesting were adopted across many lakes at once, regional synchrony in recruitment pulses could make aggregate catch more variable over time than the single-lake model suggests, a risk the paper does not model."],"forward_implications":["Managers targeting yield or stable catches can keep single-observation, biomass-based harvest control rules without sacrificing performance, even in spasmodically recruiting age-structured fisheries.","In size-dependent fisheries, a flexible policy that waits for strong cohorts and harvests in pulses can beat constant-exploitation and precautionary rules by roughly 30% in the simulated setting.","Monitoring programs that already measure mean fish weight should treat it as optional for yield and risk-averse objectives, but as potentially valuable for trophy or size-selective objectives.","Reinforcement-learned rules can produce non-intuitive feedback shapes, such as exploitation decreasing with biomass at high biomass, that would be hard to specify a priori in management strategy evaluation.","The near-tie among all optimized policies for yield and HARA utility suggests that, in this model, many different policy shapes are equivalent, giving managers freedom to pursue other objectives without sacrificing much utility."],"supporting_citations":[{"why":"Supplies the age-structured walleye model, parameter values, the Beverton-Holt recruitment form, and the empirical evidence of spasmodic recruitment.","marker":"Cahill et al. 2022"},{"why":"Defines spasmodic recruitment and documents historical stock patterns that motivate the 10–30 times recruitment pulses.","marker":"Caddy and Gulland 1983"},{"why":"Establishes the optimization and adaptive-management tradition and the bang-bang policy results that the learned trophy behavior is compared against.","marker":"Walters and Hilborn 1978"},{"why":"First showed pulse fishing can be optimal when all ages are harvested, the result the learned trophy policy rediscovers.","marker":"Walters 1969"},{"why":"Provides the deep reinforcement-learning methodology for conservation decisions that the paper adapts to harvest control rule design.","marker":"Lapeyrolerie et al. 2022"},{"why":"Introduces the proximal policy optimization algorithm used to train the neural-network harvest control rules.","marker":"Schulman et al. 2017"},{"why":"Frames the measurement-uncertainty and partial-observability problem that makes the POMDP treatment necessary.","marker":"Memarzadeh and Boettiger 2019"},{"why":"Provides the government precautionary-approach harvest strategy used as the default conventional precautionary policy comparator.","marker":"DFO 2006"},{"why":"Supplies the HARA risk-averse utility parameterization used as one of the three management objectives.","marker":"Collie et al. 2021"},{"why":"Provides the management strategy evaluation best-practices baseline that the policy-search approach extends.","marker":"Punt et al. 2016"}],"fun_headline_variants":["Trophy goal unlocks age data value in walleye management","Pulse harvest emerges when targeting old fish","Mean weight only aids trophy-focused harvest rules","Age data pays off only for trophy catches","RL finds pulse harvesting for trophy walleye"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the two-observation reinforcement-learned policies were trained long enough and stably enough that their performance reflects the true information value of the mean-weight observation, rather than an artifact of unfinished optimization.","fun_headline_variants_meta":{"raw":{"variants":["Trophy goal unlocks age data value in walleye management","Pulse harvest emerges when targeting old fish","Mean weight only aids trophy-focused harvest rules","Age data pays off only for trophy catches","RL finds pulse harvesting for trophy walleye"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000534,"raw_usage":{"total_tokens":2608,"prompt_tokens":1026,"completion_tokens":1582,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":1511}},"tokens_in":642,"tokens_out":1582,"duration_ms":10319,"temperature":1.0,"reasoning_tokens":1511,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:07:27.031438+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the HARA scenario with the same 2RL architecture but with an order of magnitude more training steps and a moving-average mean-weight observation; if that policy then exceeds the constant-U policy's utility, the claim that mean weight has no value for risk-averse management would be overturned.","supporting_citations":[],"review_version":1}