{"id":"c22f09c2-4612-4f44-ba5b-c76f33fad346","arxiv_id":"2602.23053","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A modular indoor water tank with dual motion capture, a digital twin, and a space-lab link demonstrates four experiments bridging maritime and space field robotics.","lead":"Marinarium is a modular, camera-instrumented 9×5×3 m water tank with a retractable roof, a digital twin, and a connected space-robotics lab, built to test maritime and space robots on campus. Four proof-of-concept experiments show it can support robot dynamics modeling, air-water-underwater rendezvous, simulator correction, and underwater spacecraft-surrogate validation, offering a low-cost reproducible stepping stone between simulation and field deployment.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Tank-to-field/space transfer is asserted, not demonstrated; the Sec. VII surrogate comparison uses feedback control and cannot establish dynamic equivalence.","rationale":"The reader's weakest assumption is exactly the load-bearing point: the representativeness of the tank physics, especially for microgravity surrogacy. The paper itself flags this in Sec. VII.A and Sec. VII.C, so the concern is internally acknowledged rather than an external standard being imposed. The strongest counterargument is that the facility need only be a 'testbed' for developing and debugging algorithms, not a high-fidelity physical analog. But the abstract and conclusion make a stronger claim: the facility 'bridges laboratory robotics, offshore operations, and space applications.' A debugging arena can exist without transfer validity; the claimed bridge requires it. The space-validation experiment is the clearest place this matters, because the authors explicitly compare an underwater vehicle to a planar space emulator and conclude validity. The comparison is closed-loop, so it cannot distinguish controller robustness from plant equivalence. A concrete, inexpensive open-loop test would settle this. If the test shows large hydrodynamic damping on the maneuver timescale, the central claim should be weakened; if not, the surrogate argument is substantially strengthened. The paper has real independent support: the SYSID study ships open-source code and held-out data, the facility cost estimate is specific, and the multi-domain demonstration is a genuine capability proof. These support the facility as a research infrastructure, but they do not yet support the transferability inference. Thus the reader's CONDITIONAL verdict is appropriate; my concern confirms it rather than moving it.","tokens_in":23216,"tokens_out":3499,"duration_ms":44946,"concrete_test":"Run an open-loop step-response test with the BlueROV2 in the Marinarium: command a constant wrench (as in Eq. 19) with the disturbance-estimation EKF disabled, and compare the resulting velocity profile to the drag-free spacecraft model of Sec. VII.B. Specifically, measure the hydrodynamic damping time constant tau = m/|D| from the identified dynamics (Sec. IV data). If the vehicle's velocity decays substantially on the ~20 s maneuver timescale of Figs. 11-12, then hydrodynamic forces dominate and the closed-loop tracking agreement in Fig. 13 is not evidence of microgravity equivalence. If instead the velocity stays near the drag-free prediction for the full maneuver, the surrogate claim gains quantitative support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Marinarium is an intermediate testbed between simulation and field/space deployment. What must be true is that the 9x5x3 m indoor tank is representative enough of open-water and microgravity conditions that results transfer. The paper explicitly disclaims this: Sec. VII.A states 'hydrodynamic effects introduce the main deviations from true microgravity,' and Sec. VII.C lists 'formal guarantees for equivalence between underwater and space-domain validation' as future research. This is not an external disagreement; it is the paper's own limitation statement, and it bears directly on the strongest claim.\n\nThe concrete gap is in the space-validation experiment. The BlueROV2 and ATMOS free-flyer execute the same planar STL-based inspection task under the same NMPC controller. The reported similarity in tracking error (Fig. 13) does not establish that the underwater vehicle is a valid surrogate for a free-flying spacecraft, because closed-loop feedback compensates for plant differences. The controller explicitly estimates and counteracts disturbance forces/torques (Sec. VII.B), and the residual-learning study in Sec. VI.F finds no significant improvement in angular velocity and instability on sharp turns. Thus the closed-loop agreement could be entirely due to feedback masking hydrodynamic damping, added mass, and tether forces. The same logical issue applies to the maritime side: none of the four studies compares tank behavior to offshore behavior, so the 'intermediate between simulation and field deployment' claim is supported only by plausibility, not measurement.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the Marinarium, a small (9×5×3 m) modular indoor water tank facility at KTH with underwater and above-water motion capture, a retractable roof, a digital twin in SMaRCSim, and integration with the ATMOS planar space-robotics testbed. The authors argue that this combination provides a cost-effective, reproducible intermediate step between simulation and offshore/space deployment. Four studies are used as validation: data-driven system identification of a BlueROV2 using Koopman EDMDc-RBF; a multi-domain rendezvous mission with an AUV, USV, and UAV; a learned residual-dynamics model for sim-to-real transfer; and a paired inspection task on a BlueROV2 and the ATMOS free-flyer under the same STL-based NMPC stack. The central claim is that these experiments demonstrate the facility's value for reproducible, instrumented experimentation in maritime and space-analog field robotics.","tokens_in":23523,"tokens_out":3940,"duration_ms":44149,"significance":"If the facility descriptions and experimental results are taken at face value, the paper is a useful infrastructure contribution. It is one of very few small-scale facilities with dual underwater/above-water MoCap, and the open-sourced SYSID datasets, code, and hyperparameters are a strength that increases reproducibility. The Koopman-based identification result is a plausible and interesting first application to underwater vehicles. The sim-to-real residual pipeline, despite its limitations, is a concrete demonstration of how an instrumented tank plus a digital twin can be used for underwater real-to-sim learning. The paired ATMOS/BlueROV2 experiment is also a valuable demonstration of infrastructure interoperability. However, the paper's strongest claims—that the facility is a validated intermediate testbed for offshore and space deployment—are not established by the experiments. The four studies are capability demonstrations, not transfer validations. The conclusion that the facility 'bridges laboratory robotics, offshore operations, and space applications' is substantially ahead of what the data support, and the paper's own caveats in Sec. VII.A and VII.C undercut this conclusi","major_comments":[{"comment":"The claim that the BlueROV2 'can reproduce the behavior of a free-flying spacecraft' is not supported by the closed-loop tracking comparison. Both platforms use the same NMPC controller, and the controller includes an EKF that actively estimates and compensates for disturbance forces and torques (Sec. VII.B). Similar closed-loop tracking errors can therefore be obtained even if the plant dynamics differ substantially. The paper itself acknowledges in Sec. VII.A that hydrodynamic effects are the main deviations from microgravity and in Sec. VII.C that formal equivalence is future work. The experiment demonstrates that the same software stack runs on both platforms, but not that the underwater vehicle is a valid surrogate. Either remove the validation language and reframe the result as a feasibility/prototyping demonstration, or add an open-loop dynamic comparison, a disturbance-estimate c","section":"§VII.C and Fig. 13"},{"comment":"The abstract and conclusion state that the residual model 'significantly improves simulation' of the AUV, but Table 4 shows no significant reduction in angular velocity error (e.g., at H=100: 0.23 vs 0.25; at H=500: 0.25 vs 0.27), and the text reports that the corrected model becomes unstable on sharp turns with high angular velocity and fast forward acceleration. The improvements are limited to longer horizons for pose and linear velocity. This is a load-bearing point for the sim2real bridging claim. The paper should present angular-velocity results as a negative or inconclusive result, quantify the instability (e.g., fraction of trajectories diverging, time-to-instability), and temper the abstract/conclusion accordingly. As written, the claim overstates what the experiment shows.","section":"§VI.F, Table 4"},{"comment":"The RMSE comparisons in Table 2 lack error bars, confidence intervals, or multiple train/test splits. The test set is a single chronological split, and the Koopman hyperparameters (K, gamma, lambda) are tuned on a subset of the training data, but no variance is reported. The differences between Koopman and the double-integrator baseline at 1-step and 10-step horizons (0.0629 vs 0.0784 and 0.0831 vs 0.1088) may be within experimental noise, and the paper's own Fig. 6 shows only one rollout. In addition, the PINc baseline 'did not seem to reduce significantly during training' despite using the same hyperparameters as [48], yet it is still reported as a baseline with RMSE ~8.8. This unexplained failure undermines the fairness of the comparison. Please provide multiple seeds/splits with error bars, report the PINc training issue as a failure mode rather than a silent baseline, and either fix","section":"§IV.D, Table 2, §IV.E"},{"comment":"The central claim that the Marinarium 'enables new experimental methodologies that bridge laboratory robotics, offshore operations, and space applications' is not supported by any comparison to offshore or open-water data. None of the four studies validates the tank as a representation of offshore conditions; the multi-domain mission is a demonstration of integration, not of transferability. The paper should explicitly state that the transferability of results from the tank to offshore and space environments remains an open hypothesis, and use language such as 'controlled prototyping environment' rather than 'bridge' and 'validated testbed' throughout the abstract and conclusion.","section":"§VIII, Abstract"}],"minor_comments":[{"comment":"The notation is confusing: N=9165 is called 'the total number of data points' in Eq. (13), but the dataset was earlier described as having more than 45823 samples. Please clarify whether N is the test-set size or the number of rollout starting points, and keep the definitions consistent.","section":"§IV.D, Eq. (13)"},{"comment":"The column headers (e.g., 'hat p_sim', 'p_sim', 'hat omega_sim') are hard to read with the hat notation in the table. Use separate columns with explicit 'corrected' and 'uncorrected' labels, and state the units of each quantity. Also, '7198 trajectories' appears where the surrounding text says '7998'; please check.","section":"Table 4"},{"comment":"The caption says the figure shows 'ground truth vs. Koopman, physics and DI rollouts', but the PINc model is not plotted. State this in the caption and explain why (the PINc model is left out because of poor performance).","section":"Fig. 6"},{"comment":"The order of the four research areas in the introduction (SYSID, sim2real, multi-domain, space) does not match the section order (IV, VI, V, VII). Reorder the list or the section numbering to avoid confusion.","section":"§II.C"},{"comment":"The sentence 'has enabled a transmission rate of 70% with Delphis Succorfish modems ... from a 0% without stones' is a nice quantitative result, but it belongs in Sec. V as a motivating capability for the multi-domain experiment. Also clarify the number of trials over which the 0% and 70% rates are measured.","section":"§III.B"},{"comment":"Reference [43] is cited as a digital twin of CIRTESU, but the reference title suggests a human-robot interaction paper. Double-check that this reference indeed describes a digital twin of the facility.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is an infrastructure paper with four application vignettes. The facility description is valuable and the open-sourced dataset/code is a genuine strength. However, the paper repeatedly overclaims transferability, and the space-validation experiment in particular would need either a stronger methodology or a weaker claim. The missing error bars and the unexplained PINc failure are fixable but currently undermine the quantitative comparisons. I recommend major revision rather than rejection because the central facility description is sound and the overclaims can be addressed by rewording and by adding appropriate statistical/experimental controls."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my read of arXiv:2602.23053. It's a facility paper, not a scientific breakthrough, and that's okay. What's actually new: the specific Marinarium combination — underwater and above-water MoCap, retractable roof, digital twin in SMaRCSim, and a connected planar space robotics lab — doesn't appear in any of the facilities in their Table 1. The cost estimate (under 4M SEK for the basin) and modular design are useful for anyone planning similar infrastructure.\n\nThe strongest part is the SYSID study in Section IV: Koopman EDMDc with RBF dictionary applied to a BlueROV2, with open-source code, a held-out test set, and comparisons against a double-integrator baseline and the Fossen model from [50]. The Koopman model wins at all horizons, and the result is reproducible — that's real evidence. I also give the residual sim2real study credit for reporting that it doesn't improve angular velocity and becomes unstable on sharp turns. That's honest, and it tells you the technique isn't ready for deployment.\n\nNow the soft spots. First, the RMSE tables have no error bars. The datasets are large and the splits are chronological, so some variance estimate is needed. Second, the PINc baseline simply fails to train — loss 'did not seem to reduce' — and they leave it at that. It's a baseline, so not fatal, but it weakens the comparison. Third, the acoustic-gravel claim (0% to 70% transmission with Succorfish modems) has no methodology: no link details, no packet counts, nothing. That's a significant claim for the facility and it's under-supported. Fourth, and most important, the space-validation experiment: the BlueROV and ATMOS free-flyer run the same NMPC controller including an EKF that estimates disturbance forces. So similar tracking error under feedback doesn't establish that the underwater vehicle is a surrogate for a free-flyer. The paper itself says formal guarantees for equivalence are future work. I think the stress-test note is right to flag this, but I'd temper it: the paper's concrete claim is that identical autonomy stacks can be run in both domains, and that is demonstrated. The abstract's 'intermediate testbed between simulation and field deployment' overreaches because no tank-to-offshore comparison is offered.\n\nWho should read this: anyone building or planning a maritime/space-analog testbed, and researchers interested in data-driven SYSID for AUVs. It deserves a serious referee — the facility is real, the SYSID code ships, and the limitations are mostly fixable. My recommendation to the editor: send it to peer review, with the expectation that the authors add error bars, document the acoustic measurement, and soften the transfer claims.","headline":"A real facility with a genuinely useful open-source SYSID study and honest limitations; the 'intermediate testbed' claim is plausible but transfer to field/space remains asserted rather than shown.","tokens_in":24135,"tokens_out":2517,"would_cite":true,"duration_ms":26289,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A modular, instrumented indoor water tank can serve as a reproducible, cost-effective middle ground between simulation and field deployment for both maritime and space robotics, the paper argues, with four demonstration studies backing the","keywords":["Marinarium","field robotics infrastructure","underwater robotics","sim-to-real transfer","neutral buoyancy","multi-domain robotics","system identification","motion capture"],"falsifier":"Equip the same autonomy stack used in the paired experiment, run the identical temporal-logic inspection task on the air-bearing floor, in the tank, and in an open-water or parabolic-flight setting, and compare tracking-error distributions and success times. If the tank does not rank-order the settings the way real deployments do, the intermediate-testbed claim collapses. A cheaper check specific to the sim-to-real study: retrain the residual model on tank data and evaluate on trajectories dominated by high-yaw-rate turns; the paper already notes such instability, and showing it is systematic","tokens_in":23093,"feed_emoji":"🤿","tokens_out":6892,"duration_ms":69582,"temperature":0.7,"pith_summary":"The paper aims to establish that a compact, modular water-tank facility can fill the gap between cheap simulation and costly offshore or orbital testing. Its central claim is that the Marinarium, an instrumented 9×5×3 m indoor basin with underwater and above-water motion capture, a retractable roof, and a simulator twin, provides a testbed where maritime and space-robotics experiments are repeatable and fully observable. If true, this matters because it would let many research groups collect high-quality datasets, test multi-domain fleets, close the sim-to-real gap, and validate spacecraft autonomy without access to large neutral-buoyancy labs or sea time. The paper supports the claim with four demonstrations and reports a construction cost small enough to suggest the design could be replicated.","feed_headline":"One modular tank stands in for sea and space field tests","feed_subtitle":"Underwater robots run the same autonomy software as spacecraft, making field-scale trials cheap and repeatable.","key_machinery":"The load-bearing object is the facility itself: a prefabricated, free-standing basin with a gravel floor that scatters acoustic reflections, dual motion-capture systems covering the underwater and aerial volumes, a retractable roof for real-weather operation, and a simulator twin that ingests live vehicle states. The cross-domain space experiment works because the same autonomy stack runs on both the ROV and the planar free-flyer, so software differences are removed and only the physical environment differs. Methodologically, the dynamics result rests on a Koopman operator approximated by extended dynamic mode decomposition with a radial-basis-function dictionary: the nonlinear vehicle dynam","core_discovery":"The paper argues that a single modular facility—an instrumented 9×5×3 m indoor tank with motion capture above and below the surface, a retractable roof, and a simulator twin—can be a reproducible, low-cost middle ground between simulation and field deployment for maritime and space robotics. Four demonstrations support this: a Koopman-operator method that predicts a small ROV's state in open-loop rollouts better than physics and neural baselines; a rendezvous among underwater, surface, and aerial vehicles; a learned residual-dynamics correction that shrinks simulator endpoint error; and a paired test where a neutrally buoyant ROV and a planar free-flyer execute the same temporal-logic inspec","pith_inferences":["The paper stops short of showing that tank-learned models transfer to open water; a natural next test is to deploy the identified dynamics or residual-corrected simulator against offshore data and measure how much accuracy degrades.","The paired underwater and air-bearing setup suggests a staged validation ladder—simulator, planar free-flyer, neutrally buoyant ROV, orbital demo—where the ROV leg adds realistic disturbance forces that a nearly drag-free air-bearing floor cannot provide.","The reported instability during sharp turns in the residual-dynamics study points to a concrete fix: decouple linear and angular residual predictors or train with oversampled high-yaw-rate maneuvers, so the corrected simulator can be used to train control policies.","A gravel-scattering floor that lifts acoustic-modem rates from 0% to 70% in a small tank implies that other indoor basins could become underwater-networking testbeds with minimal retrofit."],"forward_implications":["Model learning for underwater robots can use dense real-world data instead of simulation or trial-and-error field tests: the tank yielded more than 45,000 synchronized 12-dimensional state-plus-thruster samples per dataset at 50 Hz.","Underwater simulators can be made more faithful by injecting residual dynamics learned from motion-capture data; the paper reports the corrected simulator roughly halved long-horizon position error on held-out trajectories.","Spacecraft autonomy software can be rehearsed on an underwater vehicle before air-bearing or orbital testing, since the same temporal-logic plan and nonlinear MPC tracked comparably on both platforms.","Multi-domain missions can be rehearsed indoors with full ground truth, including acoustic underwater communication, which the gravel floor made possible at a 70% transmission rate.","Because the structure is assembled from prefabricated modules with construction cost estimated under four million Swedish kronor, other groups could build equivalent facilities."],"fun_headline_variants":["One tank replaces costly sea and space field tests","Indoor tank bridges sim and sea for robot trials","Modular facility runs sea and space robot tests","Underwater tank simulates both sea and space for robots","Cost-effective tank testbed for maritime and space robotics"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole value proposition rests on the assumption that hydrodynamic behavior in a 9×5×3 m indoor tank is close enough to open-water conditions, and neutral buoyancy close enough to microgravity, for results obtained there to transfer to field and space deployments; the paper itself flags hydrodynamic effects as the main deviation from true microgravity and lists formal equivalence guarantees as future work.","fun_headline_variants_meta":{"raw":{"variants":["One tank replaces costly sea and space field tests","Indoor tank bridges sim and sea for robot trials","Modular facility runs sea and space robot tests","Underwater tank simulates both sea and space for robots","Cost-effective tank testbed for maritime and space robotics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000456,"raw_usage":{"total_tokens":2135,"prompt_tokens":759,"completion_tokens":1376,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":1301}},"tokens_in":503,"tokens_out":1376,"duration_ms":10364,"temperature":1.0,"reasoning_tokens":1301,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T20:28:48.351774+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Equip the same autonomy stack used in the paired experiment, run the identical temporal-logic inspection task on the air-bearing floor, in the tank, and in an open-water or parabolic-flight setting, and compare tracking-error distributions and success times. If the tank does not rank-order the settings the way real deployments do, the intermediate-testbed claim collapses. A cheaper check specific to the sim-to-real study: retrain the residual model on tank data and evaluate on trajectories dominated by high-yaw-rate turns; the paper already notes such instability, and showing it is systematic","supporting_citations":[],"review_version":1}