{"id":"95db1589-2a2c-4d29-ae8a-6c184f0b4787","arxiv_id":"2608.04704","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper presents a multi-sensor railway dataset with over 7 million annotations across 21 classes, available from DB InfraGO upon request.","lead":"This paper describes a new railway perception dataset with over 7 million annotations of 21 object classes from two instrumented rail vehicles. It is offered as a resource for training AI systems for automated train operation, but access is by request only.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Annotated count consistency is load-bearing: the stated 7,052,055 total does not match the class totals in Table 2, and the paper neither reports sensor-wise frame distributions nor quantitative calibration validation, so the dataset's usability claim is not yet supported.","rationale":"The reader's weakest_assumption is essentially that calibration error undermines annotation alignment. I agree that quantitative calibration validation is missing, but I do not think that is the single most load-bearing condition. A more basic condition is missing per-sensor annotation-type and frame counts. The paper claims 7,052,055 annotations, but the numbers in Table 2 are only object-class totals and do not tell us how many are 3D boxes, 2D boxes, polygons, polylines, or radar points. If the total is dominated by, say, polylines or projected radar points, then the dataset is not comparable to existing 2D/3D benchmark datasets, and the central claim of a 'comprehensive multi-sensor dataset' is less useful. The total is also not tied to a clear deduplication rule or to a per-sensor frame count, so it is difficult to verify the quantity from the paper alone. The paper explicitly states that annotations were first created in 3D lidar and then projected to other sensors, and that projected annotations were checked and reworked; this is an internal consistency. However, the absence of any quantitative error metric is a genuine gap. I would not escalate to REJECT because the paper is honest about its access limitations and includes concrete sensor configurations and class breakdowns. I would keep the verdict at CONDITIONAL: the dataset may be useful, but the paper must provide per-sensor/type counts and calibration validation before it can be assessed as a benchmark. The reader's focus on calibration is plausible and partially overlaps with my concern about quantitative validation, but the missing distribution and deduplication information is more immediate and more directly load-bearing for the headline number.","tokens_in":4984,"tokens_out":1959,"duration_ms":18232,"concrete_test":"Ask the authors for a per-sequence, per-sensor, per-annotation-type breakdown (e.g., number of 2D boxes, 3D boxes, polygons, polylines, radar points) and for calibration reprojection errors (mean/median pixel error for lidar-to-RGB and lidar-to-IR, plus object-wise IoU between projected and manually corrected boxes). If the sum of all per-sensor counts does not equal the claimed 7,052,055 after deduplication, or if typical reprojection error exceeds a few pixels on the highest-resolution camera, the headline quantity and quality claims weaken. A public release of a sample sequence with raw sensor timestamps and calibration files would also settle the issue.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the dataset contains exactly 7,052,055 high-quality annotations across 21 classes (Section 6, Table 2). Summing all 21 per-class numbers in Table 2 gives 7,052,055 (GAF 4,081,482 + BR472 2,970,573 = 7,052,055), so the arithmetic is internally consistent. My concern instead targets the two conditions that would make the stated quantity meaningful. First, the 88.2 min (5292 s) of annotated data is not tied to any per-sensor frame count; annotation totals aggregate 2D boxes, polygons, polylines, 3D boxes, and projected radar points across a multi-sensor setup with different modalities. Without a per-sensor, per-annotation-type count, the headline 7 million annotations is not comparable to existing datasets such as OSDaR23 or RailGoerl24. Second, Section 4 describes projecting 3D lidar annotations onto RGB, IR, and radar frames, and Section 5 lists raw-data quality checks (calibration, odometry) only as a preliminary step. No calibration reprojection error, no quantitative agreement between projected and manually corrected annotations, and no per-class quality scores are reported. Because the supporting evidence is formal analysis only ([formal_verification: none], [parameter_count: 0]) and the dataset is not publicly released, the strongest claim rests entirely on self-reported totals. The existence of a provisional pre-print under a 2026 arXiv number also means the evaluation is based on an unfinished version. If the annotation totals are inflated by duplicated projected annotations or by counting all box coordinates without temporal deduplication, the claim would be misleading; however, I would be reluctant to assume this without a distribution report.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a multi-sensor dataset for railway environment perception, recorded from two rail vehicles (a GAF track maintenance vehicle and a BR472 commuter train) equipped with multiple RGB cameras, IR cameras, lidars, radar, and position/acceleration sensors. The authors report 7,052,055 annotations across 21 object classes, 88.2 minutes of annotated sensor data, and describe an annotation methodology that starts with 3D bounding boxes in lidar point clouds and projects them to 2D boxes, polygons, IR images, and radar images. The paper includes a comparison to existing railway datasets, a per-class annotation breakdown, and a description of quality-control steps. The dataset is available upon request from DB InfraGO AG.","tokens_in":5320,"tokens_out":3547,"duration_ms":36359,"significance":"If the dataset is made accessible and its quality can be independently verified, it would be a valuable complement to existing railway perception datasets such as OSDaR23 and RailGoerl24, offering a larger annotation count and multi-sensor coverage. The paper is a concise description of a substantial annotation effort, and the internal arithmetic of Table 2 is consistent (the sum of the class totals equals the stated overall total of 7,052,055). The use of the RailLabel/ASAM OpenLABEL schema is also a strength. However, the current manuscript does not provide sufficient quantitative evidence for the claimed annotation quality, does not report per-sensor or per-annotation-type statistics needed for comparability, and offers no public or verifiable access to the data, all of which are central for a dataset paper.","major_comments":[{"comment":"The reported total of 7,052,055 annotations aggregates across different annotation types (2D boxes, polygons, polylines, 3D boxes, and projected radar points) and across fifteen or more sensors, but the paper does not provide any per-annotation-type, per-sensor, or per-sequence frame count. Consequently, the headline figure is not interpretable or comparable to existing datasets such as OSDaR23, which report per-sensor frame counts and annotation breakdowns. Please provide a detailed breakdown of annotation counts by type, by sensor, and by sequence, at least in a supplementary table.","section":"Section 6, Table 2"},{"comment":"The paper claims 'high-quality annotations' (Abstract) and describes a projection pipeline from 3D lidar boxes to 2D boxes, polygons, IR images, and radar images, but provides no quantitative validation. There are no calibration reprojection errors, no synchronization offsets, no IoU or other agreement metrics between projected and manually corrected annotations, and no inter-annotator agreement is reported. Section 5 mentions that 'about 5% of the data was reviewed' but does not report the outcome of that review, such as an error rate or correction rate. Because the dataset's central value rests on annotation quality, this omission is load-bearing and should be addressed with concrete numbers.","section":"Sections 4 and 5"},{"comment":"The dataset is only available 'upon request' via email, with no download mechanism, persistent identifier, license, or terms of use described. This means reviewers and potential users cannot independently verify any of the reported quantities, including sensor configuration, annotation counts, or quality. Please provide at least a representative sample subset for review, a formal data access agreement, and a persistent identifier, or clearly state why this is not possible.","section":"Sections 3 and 6"},{"comment":"The annotation process is described as starting with 3D bounding boxes in lidar point clouds, followed by projection and manual reworking, but the paper does not specify which of the 21 object classes are annotated in 3D versus 2D only, nor how the attributes mentioned in Section 4 are encoded. This information is essential for users to understand the dataset's applicability and for fair comparison with other datasets. Please include a per-class annotation-type matrix.","section":"Section 4"}],"minor_comments":[{"comment":"The captions and in-text references use 'Disribution' instead of 'Distribution'.","section":"Figures 5 and 6"},{"comment":"Adding a total row for the GAF, BR472, and overall columns would facilitate verification of the stated 7,052,055 total.","section":"Table 2"},{"comment":"Please report the number of annotated frames per sensor (or at least per sequence) in addition to the total annotated duration, since frame counts are the metric used in the comparison table in Section 2.","section":"Section 6"},{"comment":"The comparison table does not list the annotation types or sensor frame rates for the proposed dataset; adding these columns would clarify how the dataset relates to the existing ones.","section":"Section 2, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The arXiv version is dated 5 Aug 2026 while the manuscript is dated September 2025, which may indicate a delayed or provisional posting; the editor may wish to confirm the intended submission status. The paper is a descriptive dataset paper with no release of data, so the review necessarily relies on the authors' self-reported numbers; the recommended revision should focus on providing verifiable statistics and a path to access."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real dataset release, not vaporware. The 7,052,055 total matches the sum of Table 2 exactly, so the internal arithmetic is sound. The stress-test note claims a mismatch, but on reading the table the sum is correct. What is actually new is the scale and sensor count. OSDaR23 had 1534 frames across camera/lidar/radar; this gives 88.2 minutes across six RGB, three IR, six lidar, and radar on the GAF, plus a similar setup on a BR472. That is a substantial extension, and the annotation methodology (3D-first, then projection to 2D and radar) is sensible. They also use RailLabel/OpenLABEL JSON, which makes downstream use easier.\n\nThe paper does what a dataset paper should: describes the vehicles, the classes, the quality-control pipeline, and the annotation types. The class breakdown is useful, and the pie charts give a rough sense of distribution. I believe the dataset exists as described.\n\nWhere it falls short is the usual gated-dataset problem. No data is released, so nobody can check annotation quality or calibration. The paper says calibration and odometry were checked, but gives no reprojection errors or sample images with overlays. There are also no per-sensor frame counts, so the 7M annotation number is not directly comparable to datasets that report per-frame boxes. That limits the headline number's meaning. Also, there is no inter-annotator agreement, but for a dataset paper that is a common omission, not a fatal one.\n\nThe dataset is not public; it is 'available upon request' from DB InfraGO. That is a real limitation for a benchmark, but it does not invalidate the paper as a resource description.\n\nOverall: this is a competent dataset paper. It deserves peer review, and I would send it to a venue like IEEE T-ITS or a dataset track. The referee should ask for (a) at least a sample of data or a public preview, (b) per-sensor and per-annotation-type counts, and (c) quantitative calibration validation. But I would not desk-reject it.","headline":"A competent dataset paper whose headline number checks out; the real limitations are access and missing quantitative quality metrics, not internal arithmetic.","tokens_in":5870,"tokens_out":2429,"would_cite":true,"duration_ms":26716,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new multi-sensor railway dataset packs 7,052,055 annotations, available on request.","keywords":["railway perception","multi-sensor dataset","grade of automation","lidar annotation","radar","object detection","environment monitoring"],"falsifier":"Open the requested dataset and compute the reprojection error: take a sample of lidar-annotated 3D boxes, project them into the camera images, and compare with manually drawn 2D boxes on those images; if the average intersection-over-union falls below a pre-specified threshold (for example, 0.5) or systematic offsets of several pixels appear, the claimed multi-sensor alignment is falsified.","tokens_in":4846,"feed_emoji":"🚆","tokens_out":10959,"duration_ms":94667,"temperature":0.7,"pith_summary":"Railway automation at every Grade of Automation—from GoA2 with a driver still on board to fully driverless GoA4—requires perception systems that detect and classify hazards in real time, and those systems need large, accurately annotated datasets. This paper reports the completion of a new multi-sensor dataset built for that purpose: over seven million annotations across 21 object classes, recorded on two rail vehicles with front-mounted RGB, infrared, lidar, and radar sensors under a variety of operational scenarios. The annotation strategy is to mark objects in the 3D lidar point clouds first and then project those labels into the camera and radar images, which gives the dataset a common multi-sensor ground truth. The authors state the dataset is now available on request from the rail infrastructure operator, positioning it as a resource for developing environment monitoring for GoA2–GoA4 systems.","feed_headline":"7,052,055 rail annotations now available on request","feed_subtitle":"Two sensor-laden trains carry RGB, IR, lidar, and radar views for 21 object classes, supporting GoA2–GoA4 automation.","key_machinery":"The mechanism that carries the dataset is the 3D-first annotation and projection pipeline. Annotators draw 3D bounding boxes, polygons, and polylines in lidar point clouds; projection functions then transfer those annotations into 2D bounding boxes on RGB camera images, 2D polygons, 2D boxes on IR images, and lidar points mapped into radar images. This makes the lidar point cloud the single source of geometric truth, and the resulting labels are exported as JSON files following a rail-specific annotation schema that builds on an open labeling standard.","core_discovery":"The central claim is the existence and practical availability of a railway perception dataset with 7,052,055 annotations across 21 classes, produced from 88.2 minutes (5,292 seconds) of annotated sensor data: 1,981 seconds in 69 sequences from a track maintenance vehicle and 3,311 seconds in 194 sequences from a commuter train. The sensor configurations include six RGB cameras, three IR cameras, six lidars, and a radar on the maintenance vehicle, and three RGB cameras, one IR camera, six lidars, and four radars on the commuter train. All objects were first annotated by experts in the 3D lidar point clouds, then projected onto the 2D camera and radar frames using projection functions, with manual checking and rework afterward. The annotation set is dominated by railway-specific classes such as catenary poles, signals, signal poles, switches, tracks, and trains, in addition to general classes like persons, road vehicles, and bicycles. The authors present this as a resource for GoA2–GoA4 automation, infrastructure monitoring, and environment observation.","pith_inferences":["The authors do not report a quantitative evaluation of projection accuracy; measuring how far lidar-derived 3D boxes deviate from hand-drawn 2D boxes would turn the qualitative statement of alignment into a number users can trust.","The heavy class imbalance—over a million catenary-pole annotations versus a few thousand bicycles or wheelchairs—means naive training on the full set will favor frequent classes; benchmark designers would need to define balanced evaluation splits.","Because the dataset contains only 88.2 minutes of annotated sequences, conclusions about generalization across seasons, weather, and geographical regions would need to be established by additional collections or domain adaptation.","If the projection pipeline is as reliable as claimed, a camera-only perception system could be trained in part using lidar-derived labels, effectively transferring 3D geometry into 2D detectors."],"forward_implications":["The dataset can be used as training and validation data for perception models that support partially automated (GoA2) through fully automated (GoA4) railway operation.","Because the annotations live first in 3D lidar space and are projected to all other sensors, the same object has aligned labels in RGB, IR, and radar, enabling multi-sensor fusion and cross-modal learning.","The class distribution, with railway infrastructure elements such as catenary poles, signals, switches, and tracks represented in the millions, covers object categories that automotive datasets largely ignore.","As an available-on-request resource, the dataset gives industry teams a common benchmark for comparing environment-monitoring algorithms without repeating the expensive data collection.","The 88.2 minutes of annotated data provide a real-world complement to synthetic railway datasets, supporting validation of simulation-trained models."],"supporting_citations":[{"why":"It supplies the earlier open multi-sensor dataset whose sensor setup and recording approach this work extends for the maintenance vehicle.","marker":"[10]"},{"why":"It documents the project that equipped the commuter train with its RGB, IR, lidar, and radar configuration.","marker":"[15]"},{"why":"It defines the JSON schema used to export the annotations in this dataset.","marker":"[16]"},{"why":"It provides the open labeling standard of which the annotation schema is a subschema.","marker":"[17]"},{"why":"It establishes the railway scene-annotation task with semantic labels for tracks, switches, and signals that this dataset expands to multi-sensor data.","marker":"[4]"},{"why":"It offers a synthetic railway dataset with pixel-level annotations, serving as a comparison point for real recorded multi-sensor data.","marker":"[12]"}],"fun_headline_variants":["7M rail annotations from dual-sensor trains","Multi-sensor rail dataset: 7M objects, 21 classes","Railway AI data: 7M annotations on request","Two trains, 7M labels for rail automation","GoA2-GoA4 dataset: 7M multi-sensor annotations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's usefulness rests on the assumption that all the sensors on each vehicle are precisely aligned and synchronized, so that the lidar-made 3D labels land on the correct pixels in the camera images and radar frames.","fun_headline_variants_meta":{"raw":{"variants":["7M rail annotations from dual-sensor trains","Multi-sensor rail dataset: 7M objects, 21 classes","Railway AI data: 7M annotations on request","Two trains, 7M labels for rail automation","GoA2-GoA4 dataset: 7M multi-sensor annotations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1345,"prompt_tokens":933,"completion_tokens":412,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":326}},"tokens_in":549,"tokens_out":412,"duration_ms":4779,"temperature":1.0,"reasoning_tokens":326,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:28:47.120586+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Open the requested dataset and compute the reprojection error: take a sample of lidar-annotated 3D boxes, project them into the camera images, and compare with manually drawn 2D boxes on those images; if the average intersection-over-union falls below a pre-specified threshold (for example, 0.5) or systematic offsets of several pixels appear, the claimed multi-sensor alignment is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the earlier open multi-sensor dataset whose sensor setup and recording approach this work extends for the maintenance vehicle."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It documents the project that equipped the commuter train with its RGB, IR, lidar, and radar configuration."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It defines the JSON schema used to export the annotations in this dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the open labeling standard of which the annotation schema is a subschema."},{"cited_title":"1221-1229","cited_arxiv_id":null,"evidence_quote":"It establishes the railway scene-annotation task with semantic labels for tracks, switches, and signals that this dataset expands to multi-sensor data."},{"cited_title":"D’Amico, F","cited_arxiv_id":null,"evidence_quote":"It offers a synthetic railway dataset with pixel-level annotations, serving as a comparison point for real recorded multi-sensor data."}],"review_version":1}