{"id":"dc5168b0-036f-41ff-828f-1e27d1072892","arxiv_id":"2509.09336","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A joint spatio-temporal model integrates survey and commercial fishery data, separating presence from biomass and separately estimating how much fishing effort tracks each one.","lead":"The paper builds a six-layer statistical model that fuses two kinds of fishery data, systematic surveys and commercial catch logs, while correcting for the fact that fishermen preferentially fish where fish are abundant. If it works, managers get more reliable sardine distribution maps from a model that uses every available data source.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Identifiability of shared field W: fixed coefficient 1 in Eqs (2)-(3) and omission from FDD intensity (5) can distort estimated preferential-sampling coefficients.","rationale":"The reader's weakest assumption identifies the identical shared coefficient for W as the key vulnerability; that is the core of my concern. I extend it by noting that W is also omitted from the FDD intensity equation, making the identifiability problem sharper for time-varying preferential sampling. The paper's own Section 5 acknowledges compensation between alpha'' and beta, which is direct evidence that the attribution among the intercept and preferential terms is not clean. This concern is concrete and testable: the current simulation study is generated under the same model structure, so it cannot reveal the failure mode. However, the paper is transparent about limitations and the model may still work under the assumed structure; the reader's CONDITIONAL verdict already reflects this. I therefore recommend no change to the verdict, with the concrete alternative-DGP simulation as the decisive check.","tokens_in":22067,"tokens_out":5022,"duration_ms":65359,"concrete_test":"Simulate under an alternative DGP in which W enters Eqs. (2) and (3) with different coefficients (e.g., c_Z = 1, c_Y = 0.5), or in which the FDD intensity is log(lambda) = alpha''(t) + beta'(t)V + beta(t)U + gamma W with gamma != 0. Fit the proposed model to 100 replicates under Scenario 3 at Comb(100,200). If the median relative bias of beta and beta' exceeds the 30-50% already reported, or if the preferential parameters are pulled toward zero, the PS interpretation is not robust. If the bias remains within the paper's current simulation bands, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is structural: W(x,t) enters both linear predictors with coefficient fixed at 1 (Eqs. 2 and 3), and the FDD intensity in Eq. (5) is built only from the static spatial fields U and V, omitting W entirely. Consequently, any time-varying shared driver is forced to shift logit(pi) and log(mu) by exactly the same amount, and the sampling intensity is assumed not to respond to the time-varying shared component of abundance. If in reality the shared dynamics affect presence and biomass at different scales, or if fishers respond to the current year's W, the model cannot represent this. The excess signal is then absorbed by U, V, and the intercepts, and the estimated beta(t), beta'(t) become composites rather than the preferential-sampling effects the paper interprets. The paper's own Section 5 documents a symptom: alpha''(t) is overestimated whenever beta(t) is underestimated, indicating weak identifiability between baseline intensity and preferential terms. Because all simulation replicates are generated from the fitted model class (with c_Z = c_Y = 1 and W absent from lambda), the simulation study cannot detect this misspecification. This directly threatens the central claim that the model can quantify preferential sampling signals and integrate FID/FDD without bias.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a six-layer hierarchical spatio-temporal model that jointly analyzes fishery-independent (FID) and fishery-dependent (FDD) data under preferential sampling (PS). The model separates a binary presence–absence process and a positive biomass process, shares a spatio-temporal latent field W between their linear predictors, adds source- and vessel-specific catchability, and models FDD locations as an inhomogeneous Poisson process whose intensity depends on two static spatial fields U and V through time-varying coefficients beta(t) and beta'(t). Inference is performed via Laplace approximation in TMB. The authors report a simulation study across three PS scenarios and four sample-size combinations, claiming accurate parameter estimation and improved prediction over single-source models, and apply the model to sardine data off southern Portugal from 2013–2018.","tokens_in":22472,"tokens_out":6333,"duration_ms":68988,"significance":"If the central claims held, the model would be a valuable contribution to fisheries stock-assessment practice: it offers a single framework for zero-inflated spatio-temporal data with distinct observation processes, explicit preferential-sampling coefficients, and vessel catchability. The Appendix likelihoods are carefully written, the TMB implementation is a sensible computational choice, and the sardine application is relevant. However, the manuscript's two load-bearing claims — that the model estimates PS effects accurately and that it can integrate FID/FDD without bias — are undermined by the simulation tables and by an untested structural identifiability assumption on the shared field W. These issues are fixable, but they are not merely presentational.","major_comments":[{"comment":"The shared spatio-temporal field W enters the presence and biomass linear predictors with the same fixed coefficient 1, and the FDD intensity depends only on the static spatial fields U and V, not on W. This structural choice forces any omitted time-varying shared driver to shift logit(pi) and log(mu) by exactly the same amount, and assumes fishing intensity does not respond to the time-varying shared component of abundance. If the true shared dynamics act on presence and biomass at different scales, or if fishers respond to the current year's W, the excess variation is absorbed by U, V, and intercepts, and the estimated beta(t), beta'(t) become composites rather than preferential-sampling effects. The simulation study is generated from the same constrained model, so it cannot detect this misspecification. I recommend an identifiability analysis and/or a misspecification simulation (e.g.","section":"Section 2.1.1, Eqs. (2)–(3); Section 2.1.3, Eq. (5)"},{"comment":"The paper's claim of 'robust performance' and 'highly accurate parameter estimates' for the preferential parameters is not supported by Table 1. In Scenario 3, the median relative bias of beta'(t) is about -0.96 to -1.03 for Comb(100,200) and Comb(100,500), meaning the presence-preferential signal is essentially estimated as zero; many 90% intervals include -1. In Scenarios 1 and 3, beta(t) is systematically underestimated by 27–50% (e.g., beta(t1) median relative bias -0.27 to -0.58). These are not 'slight biases' as stated in Section 3.4.1 and Section 5; they directly affect the model's ability to detect and quantify preferential signals, which is a central advertised contribution.","section":"Section 3.4.1, Table 1"},{"comment":"The paper acknowledges that alpha''(t) tends to be overestimated whenever beta(t) is underestimated, and describes this as a compensatory mechanism that preserves spatial predictions. This is a symptom of weak identifiability between the baseline intensity and the preferential-sampling parameters. Since the annual beta(t) and beta'(t) estimates (Figure 8) are presented as management-relevant outputs, the authors should quantify this identifiability, for example with profile likelihoods or sensitivity analyses, rather than treating the compensation as benign. Without this, the claim that the model can infer fishing behavior 'quantitatively' is not established.","section":"Section 5, Discussion"}],"minor_comments":[{"comment":"The stated median identity F_S^{-1}(0.5) = E[Z] F_Y^{-1}(0.5) is not correct for a zero-inflated distribution: if P(Z=0)=1-pi, then the median of S is 0 when pi<0.5 and otherwise equals F_Y^{-1}((pi-0.5)/pi). This identity is not used elsewhere in the paper, so it is a presentation issue, but it should be corrected.","section":"Section 2.1.1, Eq. (1)"},{"comment":"The text describing beta' underestimation as 'slight' should be reconciled with Table 1, where median relative biases close to -1 occur in Scenario 3 for n_D > 100. The language should match the reported magnitudes.","section":"Section 3.4.1"},{"comment":"The Discussion refers to 'the significant influence of ocean current direction on sardine biomass,' but the model expressions in Table 4 and the effects in Figure 7 are for current intensity (INT), not direction. Either include the direction covariate in the analysis or correct the wording.","section":"Section 4.3 / Discussion"},{"comment":"The notation for the intensity intercept is inconsistent: alpha''(t) in Eq. (5) but alpha''_t in the discussion and Supplement. Please unify.","section":"General notation"}],"recommendation":"major_revision","confidential_remarks":"The paper builds heavily on the authors' prior work (Silva et al. 2024, 2025a, 2025b; Silva and Menezes 2025), and the novelty is incremental. More importantly, the simulation evidence for the central claim is weaker than the text suggests, and the structural identifiability issue around W is not addressed. I see a path to revision: add a misspecification experiment and a more honest reporting of bias, but the current version should not be accepted without these. The journal may also want to consider whether the application adds enough beyond the authors' earlier sardine analyses."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know this paper before you skim it: the model is real and the likelihood is correctly set up, but the abstract promises more than the tables deliver, and there is a structural identifiability issue the authors notice in passing without wrestling with it.\n\nWhat's actually new: a shared AR(1) spatio-temporal field W enters both the presence and biomass linear predictors, the preferential-sampling coefficients beta(t), beta'(t) are time-varying and separated by process, and a catchability layer handles vessel effects. That combination is not in any single earlier reference, even though the building blocks are mostly the authors' own prior papers. The likelihood checks out: the Gamma density in A.1 matches the stated mean and variance, the IPP intensity in (5) is standard, and the factorization in (8) is Diggle-style and correct. The simulation study is honest in design—100 replicates, multiple sample-size combinations—and the shared-field parameters (phi_W, sigma_W, delta) and the regression coefficients are recovered well. That is genuine evidence the machinery works.\n\nThe soft spots, in order:\n\n1. The abstract says \"accuracy in parameter estimation\" and \"ability to detect preferential signals.\" Table 1 tells a different story. In Scenario 1, beta(t) is underestimated by 27–50% across every configuration; in Scenario 3, beta'(t) sometimes collapses to near zero with enormous confidence intervals, and beta(t) is off by 30–58%. The discussion concedes \"slight biases.\" That is not slight. The claims need toning down.\n\n2. The shared-field identifiability is the load-bearing assumption. W enters (2) and (3) with coefficient fixed at 1, and W is absent from the FDD intensity (5). So any time-varying shared driver is forced to shift logit(pi) and log(mu) by the same amount, and fishing intensity is assumed not to respond to the current year's W. If real fishers respond to abundance anomalies or the two processes scale differently with the shared field, the excess signal goes into U, V, and the intercepts, and your beta estimates become composites. The paper's own Section 5 shows alpha''(t) is overestimated when beta(t) is underestimated—exactly the compensation you'd expect under weak identifiability. Because every simulation replicate is generated from the fitted model class, the simulation cannot detect this. This is not fatal to the paper as a methods contribution, but it needs a misspecification experiment with W excluded from lambda or with different coefficients for the two processes.\n\n3. Eq (1) states a median identity, F_S^{-1}(0.5) = E[Z] F_Y^{-1}(0.5), that is false in general—for a zero-inflated positive variable, the median is often 0. It is inherited from the authors' 2024 paper, but it is wrong and should be corrected rather than repeated.\n\n4. Minor: no code or data shipped, and the catchability layer is never tested in simulation; in the application the AIC comparison chooses the simplest version.\n\nBottom line: this is a serious methodological paper for fisheries statisticians and anyone integrating opportunistic and designed data. It deserves a real referee, but the authors need to fix the median identity, temper the abstract, and run a misspecification check before I'd trust the beta estimates as quantitative statements about fishing behavior.\n\nRecommendation: send it to peer review. It is not a desk reject. If it lands on your desk, accept the assignment but ask for a robustness simulation.","headline":"Coherent six-layer ZI spatio-temporal model with separate PS terms for presence and biomass; worth reviewing, but the abstract oversells accuracy and the shared-W identifiability needs a misspecification check.","tokens_in":23000,"tokens_out":2911,"would_cite":false,"duration_ms":26717,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M30","62P12"],"pacs":[],"model":"deepseek-v4-flash","headline":"A six-layer joint model integrates survey and commercial fishery data under preferential sampling, separating presence from biomass to correct bias.","keywords":["data integration","preferential sampling","zero-inflated models","spatio-temporal modeling","species distribution model","fishery-dependent data","fishery-independent data","sardine"],"falsifier":"Simulate a scenario in which the shared spatio-temporal field enters the presence and biomass predictors with different coefficients, fit the proposed model with the restricted equal-coefficient structure, and quantify the bias in β(t) and β'(t); the paper's reported compensation between α''(t) and β(t) already hints that this restricted structure is where the model would lose the preferential signal.","tokens_in":21868,"feed_emoji":"🐟","tokens_out":4211,"duration_ms":46228,"temperature":0.7,"pith_summary":"The paper proposes a single spatio-temporal model that pools fishery-independent survey data (FID) with fishery-dependent commercial catch data (FDD), sources with opposite strengths: surveys are unbiased but sparse, catches are dense but collected where fishermen prefer to fish. The central claim is that the model separates the presence-absence process from the biomass process conditional on presence, and explicitly links the locations of commercial fishing to both processes through time-varying preferential-sampling coefficients, thereby correcting the bias that would arise from naively pooling the sources. Simulation experiments show accurate recovery of most parameters, detection of preferential-signal strength and sign, and better predictive distributions than using either source alone. Applied to sardine off southern Portugal, the model yields annual preferential-sampling estimates and spatio-temporal biomass maps with clearer hotspots than survey-only modeling. If correct, the framework gives fisheries managers a way to combine cheap, abundant commercial data with scarce scientific surveys without losing validity.","feed_headline":"Joint model pools survey and catch data to map sardine hotspots","feed_subtitle":"Separating presence from biomass corrects preferential sampling bias, improving distribution maps from two complementary fishery sources.","key_machinery":"The central object is a six-layer hierarchical model with factorized joint distribution L(Θ)=L(ζ,σ;y)·L(π;z)·L(λ;x)·L(U)·L(V)·L(W). Presence Z follows a Bernoulli with a logit predictor containing spatial field V and shared spatio-temporal field W; biomass Y conditional on presence is Gamma with mean from a predictor containing U and W; FDD locations follow an inhomogeneous Poisson process with log-intensity α''(t)+β'(t)V+β(t)U, whose coefficients are the time-varying preferential-sampling parameters; vessel catchability k(v) scales the relative biomass index. Inference is carried out by Laplace approximation of the marginal likelihood with automatic differentiation, using a stochastic-parti","core_discovery":"The paper's key discovery claim is that zero-inflated spatio-temporal data from two differently-sampled sources can be integrated by factorizing the joint distribution into presence-absence observations, biomass-on-presence observations, the point process of commercial fishing locations, two static spatial latent fields, one shared spatio-temporal latent field, and vessel-specific catchability effects. FDD locations are modeled as an inhomogeneous Poisson process whose log-intensity is a time-varying linear combination of the presence and biomass latent fields; the coefficients β'(t) and β(t) quantify preferential sampling for each process and each year. The shared spatio-temporal field W(x,","pith_inferences":["If the shared-field structure is identifiable in real applications, the estimated β(t) and β'(t) series could be tested against external records of fishing closures and quota changes to see whether effort shifts reflect regulation rather than fish movement; the paper does not compare its estimates to such records.","A natural robustness check is to free the coefficient of the shared spatio-temporal field in one of the linear predictors; the paper fixes it to 1 in both, and its own discussion of α''(t)/β(t) compensation suggests identifiability is the fragile point.","A seasonal or higher-order version of the AR(1) shared field might matter for small pelagics with strong within-year dynamics; the paper only treats time as annual observations.","The preferential-sampling layer could be adapted to model observer-based or port-sampling bias in other disciplines where sampling intensity depends on a latent intensity field."],"forward_implications":["Pooling FID and FDD in this framework yields better spatio-temporal distribution predictions than either source alone, particularly when FDD strongly outnumbers FID.","The annual preferential coefficients provide a quantitative window into fishing-effort allocation, potentially supporting assessments of quota effects and spatial management measures.","Separating presence and biomass layers handles zero inflation explicitly, so maps can distinguish habitat occupancy from expected abundance rather than confounding the two.","The vessel-specific catchability correction makes relative biomass indices from acoustic surveys and purse-seine catches comparable within one model.","The framework generalizes to other paired 'unbiased but sparse / dense but preferential' data sources, such as epidemiology, citizen science, and other environmental monitoring settings."],"fun_headline_variants":["Model pools survey and catch data to correct sampling bias","Zero-inflated model integrates fishery data for sardine maps","Joint spatio-temporal model fixes preferential sampling bias","Unified model maps sardines using both survey and catch data","Spatio-temporal integration of fishery data with zero-inflation"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the same unmeasured spatio-temporal signal enters the presence and biomass linear predictors with exactly the same coefficient; if real shared dynamics move occupancy and abundance on different scales, the estimated preferential-sampling signals will be distorted by compensating biases between the intercept and the preferential coefficients.","fun_headline_variants_meta":{"raw":{"variants":["Model pools survey and catch data to correct sampling bias","Zero-inflated model integrates fishery data for sardine maps","Joint spatio-temporal model fixes preferential sampling bias","Unified model maps sardines using both survey and catch data","Spatio-temporal integration of fishery data with zero-inflation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000187,"raw_usage":{"total_tokens":1165,"prompt_tokens":746,"completion_tokens":419,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":490,"completion_tokens_details":{"reasoning_tokens":338}},"tokens_in":490,"tokens_out":419,"duration_ms":4560,"temperature":1.0,"reasoning_tokens":338,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T19:17:02.589629+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a scenario in which the shared spatio-temporal field enters the presence and biomass predictors with different coefficients, fit the proposed model with the restricted equal-coefficient structure, and quantify the bias in β(t) and β'(t); the paper's reported compensation between α''(t) and β(t) already hints that this restricted structure is where the model would lose the preferential signal.","supporting_citations":[],"review_version":1}