{"id":"b4e4675c-13dc-428d-888e-eb0e87b7a995","arxiv_id":"2412.00363","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Ensemble-trained neural networks predict ship maneuvering motion and produce uncertainty estimates that grow when the vessel leaves the training data distribution.","lead":"This paper combines neural network ensembles with trajectory-based training to predict ship maneuvering motion and to flag when predictions are unreliable. It tests the approach on simulated harbor maneuvers and full-scale ship data, showing that prediction uncertainty increases when the ship moves outside the training data range.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The uncertainty claim rests on correlation, not calibration: reported Mahalanobis values suggest the ensemble is orders of magnitude overconfident, so the bias/variance trend does not establish that epistemic uncertainty is captured.","rationale":"The reader identified ensemble diversity as the weakest assumption and noted in the rationale that the uncertainty is 'not rigorously calibrated.' I agree that the ensemble is implicitly assumed to approximate the posterior over functions, but the sharper, more testable issue is that the paper's own validation metric cannot detect miscalibration. The reported Mahalanobis magnitudes are so large that, taken at face value, they indicate severe overconfidence. This does not move the verdict: the paper remains CONDITIONAL, because the qualitative correlation evidence and the PD-control worst-case results are still useful, but the central claim about epistemic uncertainty requires a calibration check before acceptance. I would not REJECT, since the authors are modest in their claims and explicitly acknowledge limitations such as inability to represent aleatoric uncertainty. I also credit the paper for testing across diverse maneuvers and full-scale data, and for comparing TS1 versus TS_infinity, which is a meaningful design choice. However, the calibration gap is load-bearing for a safety-motivated simulator, and it should be settled prior to relying on the method in control evaluation.","tokens_in":23433,"tokens_out":6195,"duration_ms":67770,"concrete_test":"Recompute, for every time step in DTest-B and DTest-ZT, the Mahalanobis distance d^2_{n,k} = (nu - mu)^T Sigma^{-1} (nu - mu) using the particle covariance of Eq. (30). Compare the time-averaged value to the chi-square expectation of 3 for three velocity degrees of freedom, and compute the empirical coverage of the 95% ellipsoid. If the mean is far above 3 or coverage is far below 95%, the ensemble is not capturing epistemic uncertainty in a calibrated sense, and the paper should either apply post-hoc calibration or weaken the claim to 'relative uncertainty indicator'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that ensemble spread captures epistemic uncertainty, but the only quantitative support is the correlation between LEucl and LMaha (Section 5.3, Eqs. 31a and 31b). Correlation is not calibration. For a calibrated three-dimensional velocity prediction, the quantity in Eq. (31b) should average approximately 3 (the mean of a chi-square distribution with 3 degrees of freedom; 15 if the authors intended squared Mahalanobis distances). The values shown in Fig. 5b appear to be on the order of 10^2 to 10^3, which implies the ensemble covariance is far too small and the predictor is severely overconfident. The observed trend 'large bias implies large variance' is consistent with a variance estimate that merely scales with local model error, but it does not show that the variance is a faithful posterior width. This matters directly for the PD-control application in Section 5.4: the worst of 100 underdispersed particles can still overestimate true-system performance, as the authors concede on the stability boundary in Section 7. Since the stated purpose is safety-relevant harbor maneuvers, the absence of any calibration or coverage test is the load-bearing weakness of the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an ensemble learning approach for non-parametric system identification of ship maneuvering motion using feedforward neural networks. A set of FNN models is trained by maximum likelihood on trajectory data, with an auxiliary FNN for initial-state estimation, and predictions are propagated as particles using the TS1 and TS∞ trajectory-sampling methods of Chua et al. The authors evaluate the method on simulated model-ship data (berthing, zigzag, turning, and random maneuvers) and on full-scale port navigation data, and they apply the ensemble simulator to heading-keeping PD control evaluation. The central claim is that the ensemble spread captures epistemic uncertainty caused by insufficient or unevenly distributed data, so that low-bias predictions show low variance and high-bias predictions show high variance, and that worst-case particle evaluation reduces overestimation of control performance.","tokens_in":23641,"tokens_out":5728,"duration_ms":61644,"significance":"If the central claim is substantiated, the proposed method would be a practically useful non-parametric maneuvering simulator for harbor operations, where data coverage is uneven and safety margins matter. The paper has several strengths: the likelihood-based training procedure is coherent, the evaluation uses an external MMG simulator as ground truth, the comparison of TS1 and TS∞ is informative, the full-scale data experiment addresses applicability, and the authors are candid about residual overestimation in the PD-control study. However, the evidence for calibrated epistemic uncertainty is currently indirect: the main quantitative support is a correlation between bias and ensemble variance, not a coverage or calibration test, and the reported Mahalanobis values indicate severe overconfidence. The safety-related claim about worst-case evaluation is also qualitative. These gaps are fixable but are load-bearing for the paper's core message.","major_comments":[{"comment":"The central claim that the ensemble captures epistemic uncertainty is supported only by the correlation between LEucl and LMaha, which is not a calibration statement. For a calibrated three-dimensional velocity prediction, the quantity in Eq. (31b) should average approximately 3 (the mean of the chi-square distribution with 3 degrees of freedom; about 3.1 with the finite-sample correction for P=100). The values in Fig. 5b are on the order of 10^2 to 10^3, implying that the ensemble covariance is far too small and that the predictor is severely overconfident. The observed 'large bias implies large variance' trend is consistent with a variance estimate that scales with local error without being a faithful posterior width. The authors should report empirical coverage of the predicted velocity ellipsoids at standard levels (e.g., 50%, 90%, 99%), quantile-quantile or reliability diagrams, and, if needed, discuss recalibration. Without such a test, the statement that the method 'captures epistemic uncertainty' is not established.","section":"§5.3, Eq. (31b), Fig. 5b"},{"comment":"The ensemble diversity premise is asserted rather than verified. The paper assumes that FNNs trained on the same dataset with random initialization and mini-batch shuffling form a plausible sample from the function space of the true maneuvering model (the model set Φ in Section 4). No metric of ensemble diversity is reported, and although the training in Section 5.3 is repeated five times, no error bars or repeatability statistics are given for the main prediction results. Deep ensembles are known to be under-diverse in general, so this assumption needs direct evidence. The authors should report the spread of LEucl and LMaha across repeated ensemble trainings, pairwise model disagreement, or another diversity measure, to support the interpretation of ensemble spread as epistemic uncertainty.","section":"§3.3, §4"},{"comment":"The claim that worst-case particle evaluation 'reduces the possibility of overestimating performance' is only qualitative. The ratio max_p L_PD,p / L_PD,true is shown on a log scale, but no quantitative statement is made about how often, or by how much, the worst-case prediction still overestimates the true system. Section 7 concedes that overestimation persists on the stability boundary even with M=75, which is exactly the regime relevant to safety-critical harbor maneuvers. I recommend quantifying the conservativeness: report the distribution of L_PD,p / L_PD,true over particles and over training seeds, state the fraction of gain combinations for which the worst of 100 particles still overestimates the true score, and discuss whether any bounded or distributionally robust guarantee can be claimed. As written, the practical safety benefit is plausible but not demonstrated at the level the introduction promises.","section":"§5.4, Figs. 9 and 15, §7"},{"comment":"The full-scale experiment is presented as evidence of applicability, but Section 7 states that the method cannot represent aleatoric uncertainty due to unobserved variables such as waves, currents, and draft changes. The failure mode shown in Fig. 13, where the bias relative to observations is larger than the particle variance over long segments, is exactly what one would expect from unmodeled aleatoric effects rather than from epistemic uncertainty about the maneuvering model. The authors should either apply the method to segments where such disturbances are small, model or filter the aleatoric component, or explicitly frame the full-scale results as a demonstration of the method's limitations. Without this, the full-scale claim is weaker than stated and the interpretability of the large Mahalanobis values is unclear.","section":"§6.2, Fig. 13, §7"}],"minor_comments":[{"comment":"There are several stylistic repetitions, e.g., two consecutive sentences beginning with 'However' in the paragraph on scale effects and full-scale measurements.","section":"Introduction, §1.1"},{"comment":"The quantity in Eq. (31b) is called the 'mean squared Mahalanobis distance,' but the formula contains no outer square; the squared distance is inside the bilinear form. Please align the terminology with the definition.","section":"Eq. (31b)"},{"comment":"The middle condition in the piecewise definition of f_step(y) reads 'ϵ < y < ϵ', which is empty as written; it should presumably be '-ϵ < y < ϵ', with the last branch 'y ≤ -ϵ'.","section":"Eq. (40)"},{"comment":"The text says 'as in Section 6.2' when referring to the evaluation procedure; this should be Section 5.3.","section":"§6.2"},{"comment":"Reference [28] is listed only as 'arXiv preprint' without an arXiv identifier; please supply the complete citation.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid application-oriented study, and the trajectory-likelihood framework is coherent. The main risk is that the central epistemic-uncertainty claim is supported by correlation rather than calibration, and the safety-oriented conclusions need quantification. These issues are addressable within the manuscript's scope, hence major revision rather than rejection. The novelty relative to the authors' own prior work ([32], [41], [56]) is incremental, so the revised version should sharpen what is new beyond the initial-state FNN and the ensemble/TS comparison."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you work on black-box maneuvering models or applied deep ensembles. The paper's contribution is narrow but real: first use of deep ensembles for epistemic uncertainty in nonparametric SI of ship maneuvering, plus an initial-state FNN that measurably improves fitting. The PD-control worst-case evaluation is sensible, and the full-scale data demonstration is a genuine plus. The authors are honest about limitations.\n\nThe soft spot is the central claim. They show that ensemble spread correlates with prediction error—large bias tends to come with larger variance, more so with TS∞. Correlation is not calibration. Their Mahalanobis metric should average around 3 for a calibrated three-dimensional velocity prediction; the plotted values are in the hundreds to thousands. That points to an ensemble that is orders of magnitude overconfident. The bias-variance trend is still useful as a qualitative indicator, but it does not support the stronger statement that the ensemble captures epistemic uncertainty. The absent empirical coverage test is the load-bearing gap.\n\nThere is also an internal inconsistency: training says the likelihood integral is solved with Euler's method, while prediction says fourth-order Runge-Kutta \"as in the training.\" One of those is wrong, and it matters for reproducibility. No code or data are provided, and the main results have no error bars despite repeated training. These are fixable.\n\nWho is this for? People building harbor-maneuver simulators or applying ensemble UQ to trajectory models. It is not a rigorous UQ paper, but it is an honest engineering demonstration. I'd send it to peer review, requiring calibration/coverage analysis and a cleanup of the integration inconsistency as conditions of acceptance.","headline":"A useful but uncalibrated application of deep ensembles to ship maneuvering SI: the correlation result is real, the uncertainty claim needs a coverage test before it can carry safety weight.","tokens_in":24152,"tokens_out":1969,"would_cite":true,"duration_ms":22540,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that an ensemble of feedforward neural networks, propagated through TS∞ particle sampling, makes ship maneuvering predictions whose spread tracks epistemic uncertainty, and that worst-case particle scores reduce…","keywords":["ship maneuvering motion","ensemble learning","epistemic uncertainty","non-parametric system identification","feedforward neural networks","trajectory sampling","probabilistic prediction","autonomous surface ships"],"falsifier":"Take the model-scale simulator and generate test trajectories that intentionally leave the berthing state distribution, such as sustained sway velocities beyond those in DTrain-B. If, over those trajectories, the ensemble's particle covariance stays small while the mean error grows, or the ground truth falls outside the predicted 95% interval at a rate far from 5%, then the spread is not tracking epistemic uncertainty. The paper's own Fig. 15 offers a partial check: with M=75, overestimation near the PD gain boundary persists, so a fully calibrated uncertainty would remain in doubt until that residual bias is explained by measurement noise rather than missing model uncertainty.","tokens_in":23232,"feed_emoji":"🚢","tokens_out":5720,"duration_ms":49501,"temperature":0.7,"pith_summary":"This paper tries to give a non-parametric ship maneuvering model an honest sense of its own ignorance. The proposed method trains many feedforward neural networks on the same trajectory data, relying only on random initialization and mini-batch shuffling for variety, and then propagates particles through randomly selected models. The claim is that the resulting spread is small where training data is dense and large where the trajectory leaves the training distribution, so the spread behaves like epistemic uncertainty. The authors test this on simulated model-ship data, training only on berthing maneuvers and predicting zigzag, turning, and random maneuvers, and on real full-scale port navigation data. They also show that using the worst predicted particle when scoring heading-keeping PD control reduces the chance of overestimating control performance relative to the true system.","feed_headline":"Ensemble ship model warns when its maneuvers are guesswork","feed_subtitle":"Trained only on berthing data, its spread stays tight where data is rich and widens where the model is guessing.","key_machinery":"The central object is the ensemble model set Φ ≡ {fθ | θ ∈ Θ} of M feedforward neural network maneuvering models, each mapping the velocity, actuator, and apparent-wind state to the acceleration vector ν̇, trained by minimizing the negative log-likelihood of the observed kinematic trajectories. The prediction mechanism is particle propagation with trajectory sampling: P particles are propagated, and each particle uses a model drawn from Φ, either once per particle (TS∞) or resampled at each time step (TS1). Because the model set is interpreted as a plausible sample from the function space of the true time-invariant maneuvering model, the covariance of the predicted particles is the epistemic uncertainty estimate. TS∞ is preferred because it does not average out predicted states across models, and the paper's empirical signature is that TS∞ shows a smaller correlation between Euclidean bias and Mahalanobis distance, meaning the variance grows when the bias grows.","core_discovery":"The paper's central claim is that an ensemble of feedforward neural networks trained by maximum likelihood on trajectory data can quantify epistemic uncertainty in ship maneuvering models. In Section 7 the authors state the key observed relationship: when the bias between the predicted particle set and the true value was small, the variance was also small, and when the bias was large, the variance tended to increase, with this trend more pronounced for the TS∞ trajectory sampling method than for TS1. The paper further claims that considering the worst-case predicted particle in a heading-keeping PD control evaluation reduces the possibility of overestimating performance relative to the true system, and that the method transfers to full-scale ship operational data.","pith_inferences":["Editorial extension: the ensemble covariance could be tested for calibration by checking whether the Mahalanobis distance of held-out true trajectories follows the expected chi-square distribution; if it does not, the spread is a relative alert rather than a calibrated probability.","The authors note that the method cannot represent aleatoric uncertainty from waves, currents, draft changes, or observation errors; a natural extension would be to add an explicit disturbance-noise term so irreducible uncertainty is separated from model-form uncertainty.","The worst-case particle rule could be turned into a control design tool, for example by using the predicted particle distribution to constrain or robustify MPC, which the paper does not attempt.","A direct comparison against Gaussian-process maneuvering models on the same berthing-only training data would clarify whether the ensemble's uncertainty estimates are competitive with kernel methods at full scale."],"forward_implications":["A simulator built this way can tell its user when it is extrapolating: berthing-trained models give tight, accurate predictions for berthing-like trajectories and wide, less accurate predictions for zigzag, turning, and random maneuvers.","Using the worst-case predicted particle rather than a single-model prediction reduces, though does not completely eliminate, overestimation of heading-keeping PD control performance relative to the true system.","Enlarging the ensemble increases the predicted variance without reducing the mean prediction bias, so ensemble size is an uncertainty-resolution parameter rather than an accuracy parameter.","The same training recipe transfers to full-scale ship operational data, so probabilistic maneuvering models can be built from routine port logs without captive model tests."],"supporting_citations":[{"why":"Supplies the maximum-likelihood trajectory-fitting method for FNN maneuvering models that the ensemble approach extends with initial-state estimation.","marker":"[32, 41]"},{"why":"Provides the deep-ensemble rationale that random initialization and shuffling make an ensemble approximate Bayesian inference for epistemic uncertainty.","marker":"[49, 50]"},{"why":"Defines the TS1 and TS∞ trajectory sampling methods compared in the paper; TS∞ is the chosen propagation method for capturing epistemic uncertainty.","marker":"[51]"},{"why":"Provide the full-scale port navigation and berthing data used to build the training and test datasets for the real-ship validation.","marker":"[1, 2]"},{"why":"Provides the MMG-based simulator, actuator response model, and wind model used as the true system for generating the model-scale training and test data.","marker":"[56]"},{"why":"Supplies the Adam optimizer used to train each ensemble member, making the random-initialization diversity mechanism operational.","marker":"[61]"}],"fun_headline_variants":["Ensemble ship model flags uncertain maneuvers","Ship motion ensemble quantifies uncertainty","Neural ensemble predicts ship motion with confidence","Probabilistic ship model warns on data gaps","Ensemble learning reveals ship maneuver blind spots"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method works only if the feedforward neural networks trained on the same dataset with random initialization and mini-batch shuffling are diverse enough that their spread behaves like a sample from the space of plausible maneuvering models; if the ensemble members are too similar or systematically miss the same region, the predicted variance will not match the actual prediction error.","fun_headline_variants_meta":{"raw":{"variants":["Ensemble ship model flags uncertain maneuvers","Ship motion ensemble quantifies uncertainty","Neural ensemble predicts ship motion with confidence","Probabilistic ship model warns on data gaps","Ensemble learning reveals ship maneuver blind spots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000441,"raw_usage":{"total_tokens":2234,"prompt_tokens":943,"completion_tokens":1291,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":1227}},"tokens_in":559,"tokens_out":1291,"duration_ms":10718,"temperature":1.0,"reasoning_tokens":1227,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:27:23.476312+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the model-scale simulator and generate test trajectories that intentionally leave the berthing state distribution, such as sustained sway velocities beyond those in DTrain-B. If, over those trajectories, the ensemble's particle covariance stays small while the mean error grows, or the ground truth falls outside the predicted 95% interval at a rate far from 5%, then the spread is not tracking epistemic uncertainty. The paper's own Fig. 15 offers a partial check: with M=75, overestimation near the PD gain boundary persists, so a fully calibrated uncertainty would remain in doubt until that residual bias is explained by measurement noise rather than missing model uncertainty.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the TS1 and TS∞ trajectory sampling methods compared in the paper; TS∞ is the chosen propagation method for capturing epistemic uncertainty."},{"cited_title":"Wakita, Model-Based Reinforcement Learning for Trajectory Tracking Control of Autonomous Surface Ship","cited_arxiv_id":null,"evidence_quote":"Provides the MMG-based simulator, actuator response model, and wind model used as the true system for generating the model-scale training and test data."},{"cited_title":"Kingma, J","cited_arxiv_id":null,"evidence_quote":"Supplies the Adam optimizer used to train each ensemble member, making the random-initialization diversity mechanism operational."}],"review_version":1}