{"id":"101a04cc-563d-4eea-8938-4582c228077a","arxiv_id":"2606.15240","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"EnvShip-Bench provides a standardized 10-minute/10-minute vessel-trajectory forecasting benchmark from Danish and U.S. AIS data with curated samples and per-sample map and neighbor context.","lead":"This paper releases EnvShip-Bench, a standardized dataset for short-term vessel trajectory forecasting built from Danish and U.S. AIS data, with 10-minute observation and prediction windows plus map and neighbor context. The submission's abstract, however, claims a larger multi-region framework with long-horizon tracks and cross-region results that the manuscript body does not contain.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Vessel-disjoint splits are asserted in the abstract but never documented in the body; without them, all ADE/FDE comparisons may be invalid.","rationale":"The reader's rejection is well-founded. I focused on the strongest load-bearing assumption: fair comparison requires vessel-disjoint train/test splits. The manuscript's top abstract explicitly promises 'vessel-disjoint splits,' but the construction section (3.3) only describes per-window caps and spacing; no split-assignment algorithm is given, and Fig. 3's geographical split maps do not prove vessel separation. Given sliding-window generation, same-vessel leakage would make benchmark scores artificially good and model rankings unreliable. This is a concrete, checkable correctness risk, not a matter of taste. The missing Greece/Norway/long-horizon content is also serious, but the split issue alone is sufficient to prevent acceptance as submitted. If the dataset is actually vessel-disjoint, the paper could be revised with documentation and code; therefore the appropriate verdict remains rejection of the current submission (unchanged from the reader).","tokens_in":10459,"tokens_out":5295,"duration_ms":50105,"concrete_test":"Download the released compact subset (or rerun the pipeline) and compute vessel-ID sets per split; check whether any vessel ID appears in more than one split, and for any such vessel, whether any selected window from the same segment appears on both sides. If the intersection is nonempty, the split is not vessel-disjoint and all baseline comparisons are compromised. If the intersection is empty, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The benchmark's central value is fair model comparison, which requires that no vessel's windows appear in both training and test. The arXiv abstract explicitly claims 'vessel-disjoint splits,' but the body never explains how splits are constructed. Section 3.3 only says the compact subset 'caps retained windows per vessel and per segment' and enforces 'minimum spacing between selected windows from the same segment' — these are intra-split redundancy controls, not cross-split vessel exclusion. Fig. 3 plots anchors by split but cannot reveal vessel-ID overlap. Because windows are generated by a sliding window over each vessel's resampled segment, any vessel shared between train and test injects near-duplicate trajectories into both; the model can memorize vessel-specific motion and inflate ADE/FDE. The conclusion that the benchmark 'is not biased toward a single model family' depends on this. This is an internal inconsistency (abstract vs. body) rather than a disagreement with community norms.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents EnvShip-Bench, a benchmark for short-term vessel trajectory forecasting built from DMA and NOAA AIS data under a 30-observation/30-prediction protocol at 20-second sampling. It consists of a large-scale core release, a quality-first compact subset (40,000 samples in Table 1), and synchronized environmental and social context extensions. Initial deterministic baselines (Seq2Seq, GRU, Bi-GRU, LSTM, Bi-LSTM, TrAISformer, Social-LSTM, Social-LSTM+Env) are reported with ADE/FDE. The body claims the benchmark provides a standardized, extensible, context-aware foundation; the arXiv abstract additionally claims a multi-region framework with 330k/106k samples and cross-region results, none of which appear in the body.","tokens_in":10732,"tokens_out":9367,"duration_ms":97373,"significance":"A well-documented, public AIS forecasting benchmark with standardized splits and aligned context would be a valuable community asset, and the effort to unify DMA/NOAA data under one protocol is commendable. The layered design (core / compact / context extensions) is sensible and the initial baseline range is appropriate. However, the manuscript currently does not establish the central validity of the benchmark: split construction is under-specified, curation thresholds are not reported, baseline results have no variance information, and the abstract promises substantially more than the body delivers. The potential usefulness of the resource does not compensate for the missing evidence.","major_comments":[{"comment":"The arXiv abstract claims a multi-region framework covering Denmark, the US, Greece, and Norway, with 330,000 short-term and 106,857 long-horizon samples, cross-region training, and scene-type-dependent context gains. The body (Sections 1, 3.1, 4.4) describes only DMA and NOAA short-term 30→30 forecasting and reports no long-horizon, no Greek/Norwegian data, and no cross-region experiments. The abstract's claim that environmental context yields largest gains in coastline-constrained scenes is not tested anywhere; Section 4.5 states that 92% of the compact subset is open-water. The manuscript must be aligned: either the missing data and experiments must be provided, or the abstract must be reduced to what the body actually delivers.","section":"Abstract vs. body"},{"comment":"The paper does not show that the train/validation/test split is vessel-disjoint. Section 3.3 caps retained windows per vessel and enforces minimum spacing between windows from the same segment, but these are intra-split overlap controls; they do not prevent the same vessel from appearing in both training and test. Section 4.3 only says all models use the same train/validation/test split. Since windows from one vessel are strongly correlated, cross-split vessel leakage would inflate ADE/FDE and make model comparisons meaningless. Please state the split construction rule and report the number of vessels per split and the number/proportion of vessels that appear in more than one split; if leakage exists, the benchmark must be rebuilt.","section":"§3.3, §4.3"},{"comment":"The pipeline description is too underspecified to be reproducible. Section 3.2 mentions conservative global motion filters, excessive temporal gaps, prolonged low-speed behavior, negligible displacement, and residual geometric anomalies without giving thresholds or formulas. Section 3.3 introduces heuristic quality and difficulty scores, caps, and minimum spacing, but does not define the scores, the cap values, the spacing value, or the stratification bins. For a benchmark paper, these are load-bearing: readers cannot verify the quality-first claim or apply the same curation. Please provide a dataset card with all thresholds, score definitions, and the resulting counts (samples, vessels, segments) per split for the core and compact releases.","section":"§3.2, §3.3"},{"comment":"All baseline numbers in Table 2 are single-run values. No random seeds, learning rates, batch sizes, model sizes, or validation-based hyperparameter choices are reported, although the abstract mentions analysis over random seeds. The conclusion that the benchmark is not biased toward a single model family depends on the relative ordering of Seq2Seq (59.57 ADE) and GRU (59.72 ADE) and TrAISformer (130.34 FDE) vs Seq2Seq (131.43 FDE); differences of 0.15 m and 1.09 m are almost certainly within run-to-run variance. Please run at least 3-5 seeds for each configuration and report mean±std, or equivalently provide confidence intervals.","section":"§4.3, Table 2"}],"minor_comments":[{"comment":"The arXiv metadata calls the system 'EnvShip,' while the full text uses 'EnvShip-Bench.' Use a single name throughout to avoid confusion.","section":"Title / naming"},{"comment":"The second paragraph begins with lowercase 'second' after a period; should be 'Second'.","section":"§4.5"},{"comment":"The 5000 m neighbor-retrieval radius is given, but there is no analysis of sensitivity to this radius or the definition of 'interaction-rich cases'; please clarify how the interaction flags are derived.","section":"§3.4.2"},{"comment":"The abstract gives a Hugging Face link and the conclusion gives a GitHub link; please unify the distribution links and include a machine-readable dataset card with checksums, schema, and license information so the release can be verified.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"This manuscript appears to combine two different paper identities: the arXiv abstract describes 'EnvShip' with multi-region/long-horizon results, while the body describes 'EnvShip-Bench' with only DMA+NOAA short-term data. The editor should ask the authors to clarify which version is under review. If the vessel-disjoint split cannot be demonstrated, or if baseline variances change the reported ranking, I would support rejection upon resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe full text of EnvShip-Bench is a reasonable short-term vessel trajectory benchmark paper, but the abstract is a different paper. The body covers a 30→30 (10 min/10 min at 20s) benchmark from DMA and NOAA, with a core set, a 40k compact subset, and aligned environmental/social context. The abstract claims multi-region Denmark/US/Greece/Norway, long-horizon tracks, cross-region experiments, and context-gain findings—none of that is in the body. That mismatch alone sinks the submission as-is.\n\nWhat is genuinely new in the body: a forecast-ready benchmark with fixed splits and per-sample context, which the cited prior AIS resources don't provide. The layered design (core, compact, context extensions) is sensible, and the limitations section is honest about the 92% open-water imbalance. If the dataset is actually released as described, it would give the maritime forecasting subfield a common testbed. That is real value.\n\nSoft spots beyond the abstract mismatch: split construction is the big one. The paper claims 'vessel-disjoint splits' in the abstract, but Section 3.3 only describes capping windows per vessel/segment and spacing within a segment—those are intra-split redundancy controls, not cross-split vessel exclusion. Nothing in the body explains how train/val/test were separated, so we can't rule out the same vessel appearing in training and test. If that happens, ADE/FDE comparisons are inflated. This needs to be documented or the claim dropped.\n\nAlso: baselines are single-run with no error bars, there's no constant-velocity or straight-line baseline, and quality/difficulty score thresholds aren't reported. These are fixable. The comparison table against existing resources is fine and fair.\n\nBottom line: the underlying asset is worth having, but the paper in its current form is not. The authors need to reconcile the abstract and body, document the splits, add uncertainty estimates and a kinematic baseline, and then it could be a solid benchmark contribution. I would not desk-reject the dataset itself—I'd send it for review with a strong request for revision. But as submitted, I'd reject.","headline":"The body is a plausible short-term maritime benchmark, but the abstract oversells multi-region, long-horizon, and cross-region results that don't exist in the text, and the vessel-disjoint split claim is undocumented.","tokens_in":11163,"tokens_out":3009,"would_cite":false,"duration_ms":29842,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"EnvShip-Bench is built to make vessel trajectory forecasting comparable: a fixed 30→30 protocol applied to public AIS data from Denmark and the US, with per-sample environmental and nearby-vessel context.","keywords":["vessel trajectory prediction","AIS","maritime benchmark","context-aware forecasting","environmental context","social context","trajectory forecasting benchmark","short-term prediction"],"falsifier":"Compute the overlap of vessel identifiers between the released training and test sets; if any vessel has windows in both, the comparability claim fails. Alternatively, train a trivial model that memorizes vessel-ID-to-output mappings and see whether its test ADE is suspiciously low.","tokens_in":10363,"feed_emoji":"🚢","tokens_out":6045,"duration_ms":62643,"temperature":0.7,"pith_summary":"This paper tries to establish a common testbed for short-term vessel trajectory prediction, arguing that the field's core problem is not missing models but missing comparability: every study uses its own preprocessing, horizons, and evaluation, so results cannot be trusted across papers. It introduces EnvShip-Bench, a benchmark built from raw AIS archives of two public sources through a single pipeline, under a uniform protocol of 10 minutes of observation and 10 minutes of prediction at 20-second resolution, in vessel-centric local metric coordinates. The release is layered—a large core set, a quality-first compact subset, and synchronized environmental and social-context extensions—so that trajectory-only, environment-aware, and interaction-aware forecasting share one split and one metric. If the benchmark is adopted, published ADE/FDE numbers would mean the same thing everywhere, and context-aware maritime forecasting could be studied systematically. The initial baselines support the paper's claim that the benchmark is neither trivial nor saturated: simple recurrent models stay competitive, while kinematic augmentation and context help only in specific configurations.","feed_headline":"Benchmark makes ship-trajectory results comparable","feed_subtitle":"A 30→30 protocol with aligned environmental context turns raw AIS data into a shared testbed.","key_machinery":"The load-bearing mechanism is the 30→30 protocol: each sample contains 30 observed and 30 future positions sampled every 20 seconds (10 minutes each), placed in vessel-centric local metric coordinates with the last observed point as origin. This single choice converts heterogeneous AIS archives into one evaluation space and makes cross-region comparability possible. Around it, the layered release does the work: the core set preserves motion diversity, the compact subset's quality scoring, stratified sampling, and per-window redundancy caps enable efficient and reproducible experiments, and the context extensions (vector/raster environmental layers plus target-centric neighbor descriptors) le","core_discovery":"The paper's central claim is that a benchmark, not a model, is what the field needs next. EnvShip-Bench organizes raw AIS data from two public archives into a standardized forecasting setting: 30 observed positions and 30 future positions at 20-second intervals, expressed in vessel-centric local metric coordinates so that Danish and US waters are directly comparable. The benchmark is released in layers: a large-scale core set, a curated 40,000-sample compact subset with stricter motion screening and redundancy control, and context extensions that align each sample with shoreline/occupancy raster and vector layers and with nearby-vessel descriptors such as relative distance, velocity, CPA, an","pith_inferences":["The fair-comparison claim would be directly testable by releasing vessel-ID overlap statistics between train and test; absent that, the split-integrity assumption remains unverified.","A scene-centric subset, which the paper flags as future work, would likely be the variant that makes environmental context matter most, since port and narrow-channel cases are exactly where shoreline structure constrains motion.","The social-context extension with CPA/TCPA descriptors could be reused beyond trajectory forecasting, for encounter-rate and collision-risk studies, because it provides per-sample encounter geometry aligned with trajectories."],"forward_implications":["If the benchmark is adopted, published ADE/FDE results become meaningful across papers because they all follow the same protocol, split, and coordinate system.","The three task variants let researchers isolate the contribution of environmental versus nearby-vessel context without changing the prediction target or the evaluation metric.","The baseline numbers give future work a concrete reference: simple recurrent models are hard to beat, and kinematic augmentation is not universally helpful.","The benchmark's documented long-tail imbalance (roughly 92% open-water, weak-interaction samples) points to rebalancing and scene-centric subsets as the next steps."],"fun_headline_variants":["EnvShip benchmark: comparable vessel forecasts across regions","Ship-trajectory forecasting gets one shared yardstick","Cross-region vessel benchmark standardizes AIS forecasting","One benchmark to compare all vessel trajectory predictions","Unified testbed for ship-motion forecasting results"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The benchmark's fair-comparison claim depends on no vessel's trajectories appearing in both training and test; the paper describes per-window and per-segment caps but never states vessel-level disjointness, so a vessel present in both splits would inflate reported errors.","fun_headline_variants_meta":{"raw":{"variants":["EnvShip benchmark: comparable vessel forecasts across regions","Ship-trajectory forecasting gets one shared yardstick","Cross-region vessel benchmark standardizes AIS forecasting","One benchmark to compare all vessel trajectory predictions","Unified testbed for ship-motion forecasting results"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00092,"raw_usage":{"total_tokens":3817,"prompt_tokens":813,"completion_tokens":3004,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":2941}},"tokens_in":557,"tokens_out":3004,"duration_ms":23608,"temperature":1.0,"reasoning_tokens":2941,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T04:40:40.539923+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the overlap of vessel identifiers between the released training and test sets; if any vessel has windows in both, the comparability claim fails. Alternatively, train a trivial model that memorizes vessel-ID-to-output mappings and see whether its test ADE is suspiciously low.","supporting_citations":[],"review_version":1}