{"id":"79cd291b-6865-44df-8e49-1cd21916aa33","arxiv_id":"2502.01946","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper introduces the first public dataset combining 4D radar, spinning radar, and FMCW LiDAR, with repeated traversals and per-sensor ground truth for SLAM and place recognition.","lead":"HeRCULES is a new driving dataset that records routes with two different radar types, a FMCW LiDAR, cameras, and GPS. It is built for testing mapping, navigation, and place recognition across sensors and repeated visits.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground truth poses from RTK-GPS/INS plus B-Spline may be inaccurate in GNSS-challenged sequences, making the per-sensor trajectories and Table IV benchmark errors unverified.","rationale":"The paper is strongest as a dataset contribution: the hardware configuration, calibration pipeline, sequence list, and comparison table are concrete and give the firstness claim good support. The reader's conditional verdict is appropriate because the actual data files are not accessible from the manuscript, so ground truth accuracy, calibration quality, and completeness cannot yet be verified. My stress-test lands on the same assumption the reader flagged: per-sensor GT derived from RTK-GPS/INS plus B-Spline interpolation is the backbone of both the SLAM and the place recognition evaluations, and it is exactly the part most likely to degrade in the advertised challenging environments. I do not see an internal inconsistency or a reason to reject the paper; the condition is on releasing validated GT and tools. Thus the verdict stays CONDITIONAL, pending the concrete validation outlined above.","tokens_in":10261,"tokens_out":3434,"duration_ms":36454,"concrete_test":"Download one GNSS-challenged sequence, e.g., Bridge 01 or Mountain 01. For every UTC time, extract the RTK/inertial status flag from the NovAtel log and compute the difference between the raw 50 Hz INS pose and the B-Spline-interpolated pose at each sensor timestamp; also transform the Aeva GT trajectory into the Navtech frame using the published extrinsics and measure the translation/rotation mismatch at shared times. If any non-fixed interval exceeds a few seconds, or if cross-sensor GT mismatch is above the expected mounting uncertainty (tens of centimeters), then the Table IV ATE/RPE figures and the relative rankings need to be re-evaluated or reported with explicit error bars.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central firstness claim about heterogeneous radar is plausible: Table I situates HeRCULES against existing datasets, and the inclusion of both the ARS548 4D radar and the Navtech RAS6 spinning radar with the Aeva FMCW LiDAR appears to be a new combination. The load-bearing weak point is the claimed 'ground truth pose for each sensor' (Sec. IV-C). That ground truth is produced by assuming the RTK-GPS/INS solution is fixed and converged before each sequence, then using B-Spline interpolation to re-time poses to each sensor's clock. This is only reliable if (i) RTK stays fixed or the INS bounds drift in every environment, including Mountain, Bridge over the Han River, Street congestion, and urban areas, and (ii) B-Spline interpolation between 50 Hz INS samples does not introduce significant error during the high dynamics the paper itself notes (speed bumps, 60 km/h bridge driving, frequent stops). No failure flags, uncertainty estimates, or independent validation such as loop-closure residual reports are given for these ground truths. Since the benchmark conclusions in Table IV, for example 4D-radar-only ATE of 64.884 m versus Fast-LIO 0.358 m, depend directly on these poses, unquantified GT error could change the ranking or exaggerate single-radar limitations. This concern is about validation, not about the sensor combination claim; it can be settled only from released data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HeRCULES, a multi-modal urban dataset combining a Continental ARS548 4D radar, a Navtech RAS6 spinning radar, an Aeva Aeries II FMCW LiDAR, stereo cameras, an IMU, and an RTK-GPS/INS. The authors claim it is the first public dataset to include both 4D radar and spinning radar alongside FMCW LiDAR. The dataset comprises 21 sequences across eight environments (Mountain, Library, Sports Complex, Parking Lot, River Island, Bridge, Street, Stream) under diverse weather and lighting conditions, with repeated visits to support place recognition and multi-session SLAM. They provide per-sensor ground truth poses derived from RTK-GPS/INS with B-Spline interpolation, and they report SLAM benchmarks (Fast-LIO, 4DRadarSLAM, ORORA) and place recognition benchmarks (Scan Context) on a subset of sequences. The paper also states that ROS tools and radar format conversion software are released.","tokens_in":10488,"tokens_out":5783,"duration_ms":56678,"significance":"If the data and tools are publicly available as stated, HeRCULES fills a genuine gap: no existing public dataset combines a 4D phased-array radar, a 360-degree spinning radar, and an FMCW LiDAR with Doppler velocity. This enables research on heterogeneous radar fusion, cross-sensor place recognition, and radar-LiDAR SLAM. The diversity of conditions, intentional revisits, and per-sensor ground truth are useful features. The authors also provide baseline evaluations, which are valuable for dataset adoption. However, the significance of the benchmark claims depends on the reliability of the ground truth poses and calibration, which are not quantitatively validated in the manuscript.","major_comments":[{"comment":"The ground truth poses are load-bearing for all benchmark conclusions, but their accuracy is not validated. The manuscript states only that the GNSS solution was fixed and the INS converged before each sequence, and that B-Spline interpolation is used to re-time poses to each sensor. No failure flags, covariance estimates, or independent checks (e.g., loop-closure residuals or comparison with an offline LiDAR SLAM) are provided for the challenging sequences such as Mountain, Bridge, Street, and urban areas where RTK-GPS can degrade. Since Table IV reports ATE/RPE values against these poses (e.g., 4DRadarSLAM ATE of 64.884 m vs. Fast-LIO 0.358 m on Sports Complex 01), any systematic GT error would directly change the numerical results and potentially the ranking. Please add per-sequence GT quality metrics, interpolation error bounds, and at least one independent validation of the GT poses.","section":"Sec. IV-C and Table IV"},{"comment":"The benchmark evaluation is too limited to support the paper's broad conclusions. Only Sports Complex 01 and Library 01 (both day, clear conditions) are used for SLAM evaluation in Table IV, and only these two are used for place recognition in Table V. The abstract, introduction, and conclusion claim that the evaluations identify limitations of single-radar SLAM in \"various environments\" and support the need for heterogeneous radar SLAM, but no night, rain, snow, bridge, mountain, or stream sequences are benchmarked. Single-run results are reported without error bars or repeated trials. Please expand the evaluation to a representative subset of adverse-weather and challenging-terrain sequences, or reframe the conclusions as preliminary case-study results rather than general findings.","section":"Sec. V-A, Table IV, Table V, and Sec. VI"},{"comment":"Quantitative extrinsic calibration accuracy is not reported. Section III-B4 presents only qualitative overlays in Fig. 4(c)-(e), and Section IV-C relies on extrinsics to produce per-sensor ground truth poses. Without quantitative metrics (e.g., reprojection error, point-to-plane residuals, or cross-sensor consistency checks), the accuracy of the per-sensor trajectories is unknown. Please report calibration residuals or an independent consistency evaluation for all sensor pairs.","section":"Sec. III-B and Sec. IV-C"}],"minor_comments":[{"comment":"The header structure of Table I is confusing, with sub-columns for \"4D Radar\" and \"Scanning Radar\" interleaved with general columns. Please restructure the table so that each sensor type is clearly separated.","section":"Table I"},{"comment":"The header row of Table II is garbled (\"FrequencyRange Azimuth Elevation Range Azimuth Elevation\") and contains the typo \"elavation\". Please reformat the table and clarify the units for each column.","section":"Table II"},{"comment":"The sentence \"The ablation study conducted for the Library shows the results for thresholds of 10 m, 15 m, and 20 min Fig. 10\" has a typo (\"20 min\" should be \"20 m\") and a missing \"in\" before \"Fig. 10\".","section":"Sec. V-B"},{"comment":"The statement \"Before logging each sequence, we ensure that the GNSS solution is fixed and the INS solution has converged\" is vague. Please specify the RTK-GPS fix status criteria and whether the status is monitored during each sequence, not only at start-up.","section":"Sec. IV-C"},{"comment":"The explanation for 4DRadarSLAM's poor performance (\"the point cloud contains fewer points than the Oculii radar\") is informal. Please report quantitative point counts per scan and mention whether any preprocessing was applied.","section":"Sec. V-A"},{"comment":"The sentence \"All the above datasets are limited to 2D radar\" at the end of Sec. II-B is ambiguous because Table I includes 4D radar datasets. Please clarify that the sentence refers only to the spinning-radar datasets discussed in that subsection.","section":"Sec. II-B"}],"recommendation":"major_revision","confidential_remarks":"The dataset appears novel and potentially valuable, and the firstness claim is plausible. The main risk is the unvalidated ground truth; this is fixable with additional analysis and should not by itself lead to rejection. I would encourage the editor to request a revised version with GT validation, expanded benchmarks, and quantitative calibration results. I did not find evidence of a novelty-disclosure problem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the headline: the paper's central claim checks out. Table I is accurate as far as I know—no prior public dataset combines a 4D radar and a mechanical scanning radar, let alone with an FMCW LiDAR. That makes HeRCULES genuinely useful for anyone working on radar-LiDAR fusion or cross-sensor place recognition. The authors also provide per-sensor ground truth poses, a calibration pipeline, ROS tooling, and revisit-rich sequences, which is more than many dataset papers offer.\n\nWhat's good: the sensor suite is well-chosen and described concretely. The calibration section is practical; using the line-index channel to handle solid-state LiDAR in the radar-camera calibration is a sensible adaptation. The benchmark evaluations, while brief, are honest—they admit the 4D radar point cloud is sparser than the Oculii used in the original 4DRadarSLAM, which explains the poor ATE. That kind of candor helps.\n\nThe soft spots are about validation, not about the sensor combination. The per-sensor ground truth is the load-bearing component, since every ATE in Table IV depends on it. The authors state they ensure the GNSS fixed and INS converged before each sequence, then B-spline-interpolate to each sensor clock. That is fine for open-sky roads, but the dataset includes a bridge over the Han River, a mountain road, and congested street traffic. No uncertainty estimates, failure flags, or independent checks like loop-closure residuals are given. So the worry that the 64.9 m ATE for 4DRadarSLAM might be partly a ground-truth artifact is real, though it doesn't undermine the firstness claim. The benchmark results also lack error bars and cover only two sequences, which is fine for a dataset paper but should not be read as a rigorous algorithm comparison.\n\nThe tone is a bit promotional—'unparalleled', 'sets a new standard'—but that's common in this genre and doesn't affect the technical content.\n\nWho benefits: radar SLAM, place recognition, and sensor-fusion researchers who need a benchmark with both radar types and a solid-state LiDAR. The paper is a dataset resource, not a theoretical contribution, so its value depends on the data actually being released and the GT being trustworthy.\n\nRecommendation: yes, send it to peer review. The firstness claim deserves scrutiny, and the GT validity should be checked against released data. If those hold, it's a solid contribution.","headline":"Useful first-of-kind sensor combination for radar fusion research, but the ground-truth poses need independent validation before trusting the benchmark numbers.","tokens_in":10971,"tokens_out":2911,"would_cite":true,"duration_ms":24664,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents HeRCULES, the first public dataset to combine a 4D phased-array radar, a 360-degree spinning radar, and an FMCW LiDAR with per-sensor ground truth, intended to support multi-session radar SLAM and cross-sensor place…","keywords":["radar SLAM","4D radar","spinning radar","FMCW LiDAR","multi-session SLAM","place recognition","urban dataset","sensor calibration"],"falsifier":"Pick a sequence with a deliberate revisit, such as Mountain 01 or Stream 01, and compute the gap between the provided ground-truth poses at the two visits to the same place; a large gap, or a disagreement larger than the reported ATE when the revisited segment is matched by a high-accuracy LiDAR map, would show the ground truth is not accurate enough to support the benchmark conclusions.","tokens_in":10079,"feed_emoji":"📡","tokens_out":8022,"duration_ms":68890,"temperature":0.7,"pith_summary":"Radar SLAM methods are usually built around one radar type: either a 360-degree spinning radar that excels at mapping and place recognition, or a lightweight phased-array 4D radar that returns range, azimuth, elevation, and Doppler velocity. This paper argues that the two types are complementary, yet no public dataset contains both, so algorithms never have to cope with the combination. HeRCULES is introduced as the first public dataset to pair a 4D radar, a spinning radar, and an FMCW LiDAR, with stereo cameras, IMU, and RTK-GPS, across eight urban environments and varied weather, lighting, and traffic. The routes repeatedly revisit the same places, and ground-truth poses are provided for every sensor individually, which makes multi-session SLAM, place recognition, and radar-LiDAR fusion directly testable. Benchmark results on the dataset show LiDAR odometry far outperforming radar-only odometry and cross-modal place recognition far weaker than same-modal recognition, framing the open problems the dataset is meant to serve.","feed_headline":"First dataset combines two radar types with FMCW LiDAR","feed_subtitle":"One benchmark for radar-LiDAR fusion, cross-sensor place recognition, and multi-session SLAM, with per-sensor ground truth.","key_machinery":"The load-bearing mechanism is the recording and calibration pipeline that makes heterogeneous sensors commensurable. The FMCW LiDAR and 4D radar are operated with a shared relative-velocity convention; the LiDAR and spinning radar are aligned through correlative scan matching on polar and Cartesian images; the LiDAR, 4D radar, and cameras are calibrated together with a reflector-based tool that uses the 4D radar's direct elevation measurement; and the IMU-LiDAR transform is initialized with a targetless method. The per-sensor ground truth is produced by taking the RTK-GPS/INS trajectory and interpolating it to each sensor's timestamps with B-Splines, so that every point cloud and image has a pose that accounts for both lever-arm and time offset. The benchmark protocols, Fast-LIO, 4DRadarSLAM, ORORA, and Scan Context, then convert the raw recordings into the quantitative claims above.","core_discovery":"The paper's central claim is that a heterogeneous radar dataset is a missing resource for radar SLAM, and that HeRCULES fills that gap: it is the first dataset to integrate a 4D radar and a spinning radar alongside an FMCW LiDAR, giving every sensor its own ground-truth pose so that trajectories and recognition results can be compared honestly. The evaluation underlines the point by showing what currently fails. On the Sports Complex and Library sequences, Fast-LIO on the FMCW LiDAR reaches an absolute trajectory error of about 10 m, ORORA on the spinning radar is less accurate, and 4DRadarSLAM on the 4D radar is far less accurate, partly because the Continental sensor returns fewer points than the Oculii radar used when 4DRadarSLAM was introduced. In place recognition with Scan Context, same-modal LiDAR queries achieve AUC near 0.97, same-modal 4D radar queries are lower, and cross-modal radar-query-on-LiDAR-database recognition falls to roughly 0.28-0.40. The paper presents these gaps as the motivation for the dataset: single-radar SLAM and direct cross-sensor matching are both open, and the data needed to work on them now exists.","pith_inferences":["If the per-sensor ground truth is as accurate as claimed, the dataset can double as a testbed for temporal synchronization: residual time offsets between sensors should appear as consistent per-sensor pose shifts on the same trajectory.","The large radar-versus-LiDAR performance gap suggests that 4D radar SLAM might benefit substantially from point densification or temporal accumulation before matching; because the raw Continental point clouds are shipped, this hypothesis is testable without new hardware.","The weak cross-modal place recognition numbers are likely a property of the geometric Scan Context descriptor rather than of the sensors themselves, so the dataset's revisits make it well suited for training or evaluating learned cross-modal descriptors.","Because the same places are scanned by three active range sensors with different fields of view and densities, the dataset can serve as a benchmark for radar point upsampling and cross-sensor completion, an application the paper mentions only briefly."],"forward_implications":["Radar-only odometry is the current weak link: in the provided sequences, 4DRadarSLAM's ATE is tens of meters while LiDAR odometry is near ten meters, so heterogeneous or fused SLAM is the obvious next target.","Cross-modal place recognition has a concrete baseline to beat: matching a 4D radar query against an FMCW LiDAR database yields AUC around 0.28-0.40 on the tested sequences, far below same-modal recognition.","Multi-session SLAM can be evaluated directly because the routes revisit locations and every sensor has its own pose, which the Mountain, River Island, and Stream sequences are designed to support.","The shared Doppler-velocity structure of FMCW LiDAR and 4D radar makes the dataset suitable for comparing velocity estimation across the two sensing principles."],"supporting_citations":[{"why":"Supplies the correlative scan-matching method used to calibrate the FMCW LiDAR to the spinning radar.","marker":"[25]"},{"why":"Provides the spinning-radar polar-image format that the dataset's conversion tools target.","marker":"[22]"},{"why":"Supplies the Scan Context descriptor used for both same-modal and cross-modal place recognition evaluation.","marker":"[37]"},{"why":"Fast-LIO is the LiDAR-inertial SLAM baseline whose trajectory error anchors the comparison.","marker":"[35]"},{"why":"4DRadarSLAM is the 4D radar SLAM baseline whose poor accuracy motivates heterogeneous and fusion SLAM.","marker":"[11]"},{"why":"ORORA is the spinning-radar odometry baseline that provides the intermediate SLAM result.","marker":"[36]"},{"why":"Provides the continuous-time B-Spline interpolation used to assign ground-truth poses to each sensor's timestamps.","marker":"[34]"},{"why":"Supplies the joint camera-LiDAR-radar calibration tool, extended here with the 4D radar's direct elevation measurement.","marker":"[32]"}],"fun_headline_variants":["First dataset with both 4D and spinning radar plus FMCW LiDAR","Heterogeneous radar meets LiDAR for multi-session SLAM","Dual-radar SLAM benchmark with FMCW LiDAR and ground truth","HeRCULES: combining two radar types with LiDAR for SLAM","Urban SLAM benchmark fuses spinning and 4D radars with LiDAR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's reliability rests on the assumption that the RTK-GPS/INS poses, interpolated to each sensor's timestamp, remain accurate in urban canyons, under bridges, on mountain roads, and in dense traffic, so that every reported error is really a sensor or algorithm error rather than a ground-truth error.","fun_headline_variants_meta":{"raw":{"variants":["First dataset with both 4D and spinning radar plus FMCW LiDAR","Heterogeneous radar meets LiDAR for multi-session SLAM","Dual-radar SLAM benchmark with FMCW LiDAR and ground truth","HeRCULES: combining two radar types with LiDAR for SLAM","Urban SLAM benchmark fuses spinning and 4D radars with LiDAR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000332,"raw_usage":{"total_tokens":1901,"prompt_tokens":1055,"completion_tokens":846,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":671,"completion_tokens_details":{"reasoning_tokens":745}},"tokens_in":671,"tokens_out":846,"duration_ms":8273,"temperature":1.0,"reasoning_tokens":745,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T13:51:53.820702+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Pick a sequence with a deliberate revisit, such as Mountain 01 or Stream 01, and compute the gap between the provided ground-truth poses at the two visits to the same place; a large gap, or a disagreement larger than the reported ATE when the revisited segment is matched by a high-accuracy LiDAR map, would show the ground truth is not accurate enough to support the benchmark conclusions.","supporting_citations":[{"cited_title":"Boreas: A multi-season autonomous driving dataset,","cited_arxiv_id":null,"evidence_quote":"Supplies the correlative scan-matching method used to calibrate the FMCW LiDAR to the spinning radar."},{"cited_title":"The oxford radar robotcar dataset: A radar extension to the oxford robotcar dataset,","cited_arxiv_id":null,"evidence_quote":"Provides the spinning-radar polar-image format that the dataset's conversion tools target."},{"cited_title":"Scan context: Egocentric spatial descrip- tor for place recognition within 3d point cloud map,","cited_arxiv_id":null,"evidence_quote":"Supplies the Scan Context descriptor used for both same-modal and cross-modal place recognition evaluation."},{"cited_title":"Fast-lio: A fast, robust lidar-inertial odometry package by tightly-coupled iterated kalman filter,","cited_arxiv_id":null,"evidence_quote":"Fast-LIO is the LiDAR-inertial SLAM baseline whose trajectory error anchors the comparison."},{"cited_title":"4dradarslam: A 4d imaging radar slam system for large-scale environments based on pose graph optimization,","cited_arxiv_id":null,"evidence_quote":"4DRadarSLAM is the 4D radar SLAM baseline whose poor accuracy motivates heterogeneous and fusion SLAM."},{"cited_title":"Orora: Outlier-robust radar odometry,","cited_arxiv_id":null,"evidence_quote":"ORORA is the spinning-radar odometry baseline that provides the intermediate SLAM result."},{"cited_title":"Continuous-time visual-inertial odometry for event cameras,","cited_arxiv_id":null,"evidence_quote":"Provides the continuous-time B-Spline interpolation used to assign ground-truth poses to each sensor's timestamps."},{"cited_title":"A joint extrinsic calibration tool for radar, camera and lidar,","cited_arxiv_id":null,"evidence_quote":"Supplies the joint camera-LiDAR-radar calibration tool, extended here with the 4D radar's direct elevation measurement."}],"review_version":1}