Pith. sign in

REVIEW 2 major objections 6 minor 68 references

Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving

T0 review · 2 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper presents STU, the first publicly available dataset for road anomaly segmentation with dense 3D semantic labels, LiDAR and camera data, and temporal sequences, and shows that 2D-derived baselines struggle on it.

desk verdict First public 3D LiDAR anomaly segmentation benchmark with real value; needs OOD verification and label-quality evidence before it becomes standard. read the letter →

arxiv 2505.02148 v1 pith:YB435P2I submitted 2025-05-04 cs.CV

classification cs.CV
keywords 3DLiDARanomalysegmentationautonomousdrivingout-of-distributiondetectionpointclouddatasetbenchmarkroaddebrissemantictemporalsequences
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents STU, a new dataset for anomaly segmentation in autonomous driving, built from 128-beam LiDAR, eight cameras, and sequential recordings. Its central claim is that STU is the first publicly available benchmark with dense 3D semantic labels for road anomalies, combining LiDAR and camera data with temporal information. The authors argue that the field needs such a resource because existing anomaly segmentation research is dominated by 2D image benchmarks, while autonomous vehicles rely on LiDAR for range and robustness. Baseline experiments adapting 2D anomaly methods to 3D show that these methods perform far worse than in 2D, with models confidently labeling anomaly objects as inliers. If the dataset and its labels hold up, it supplies a missing evaluation ground for 3D and multimodal anomaly segmentation.

What carries the argument

The load-bearing object is the dataset itself: 70 fully annotated sequences with a 128-beam LiDAR, dense point-level labels for inlier, anomaly, and unlabeled classes, and per-instance anomaly IDs, plus two anomaly-free sequences used to reduce the domain gap with SemanticKITTI. The annotation pipeline starts from pseudo-labels generated by a SemanticKITTI-trained model and refines them with three annotators, and the evaluation protocol follows SemanticKITTI's 50-meter range and requires at least five points per anomaly instance. This protocol turns the raw point clouds into a reproducible testbed where point-level metrics (AUROC, FPR@95, AP) and object-level metrics (PQ, UQ) can be computed for any 3D segmentation model.

What would settle it

Independently re-annotate a random subset of the 19 validation sequences with fresh annotators who do not see the published labels, then compute agreement on point-level anomaly labels. If agreement is low, or if the fresh labels contradict the published labels on a substantial fraction of anomaly points, the reported baseline scores cannot be treated as a reliable benchmark.

Watch

Extended reading notes

Core claim

The core discovery is a public benchmark that makes 3D anomaly segmentation measurable for the first time. The dataset defines two label classes, inliers and outliers, with instance-level identity for each anomaly, and adds extra sequences without anomalies to train in-distribution models. The authors show that when standard 2D anomaly segmentation techniques—Max-Logit, Monte Carlo Dropout, Deep Ensembles, a void classifier, and RbA—are adapted to a Mask4Former-3D backbone, their point-level and object-level scores on STU are markedly lower than their 2D counterparts. Large anomalous objects are frequently predicted as familiar inlier classes such as "other vehicle" with high confidence, producing high false-positive rates at 95% recall and low average precision. The conclusion the paper draws is that directly transferring 2D anomaly methods to LiDAR does not work and the community needs 3D-specific approaches.

Load-bearing premise

The ground-truth labels are trustworthy enough to benchmark methods, even though they originate from machine pseudo-labels and the paper reports no inter-annotator agreement or label-error analysis.

Editorial extensions

If this is right

  • 3D anomaly segmentation can now be evaluated on real LiDAR data with dense labels, giving researchers a shared reference for comparing methods.
  • The poor baseline results imply that methods designed for 2D images, such as max-logit scoring or ensembling, do not transfer directly to LiDAR point clouds.
  • The dataset's temporal and multimodal setup makes it possible to design and test anomaly methods that exploit several frames or camera-LiDAR fusion.
  • Because anomalies appear as very few points among roughly 100,000 inlier points, the benchmark exposes class imbalance as a core difficulty for future methods.
  • The release of training, validation, and a closed test set with a submission procedure allows the community to track progress over time.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the pseudo-labeler that seeds the annotations is biased toward SemanticKITTI classes, the benchmark may inherit that bias; a useful extension would be measuring agreement between human annotators and comparing labels produced by different seed models.
  • The 5-point threshold for evaluating an anomaly instance excludes many of the smallest and most distant objects shown in the dataset histograms, so the benchmark likely underestimates the true difficulty of far-range anomaly detection.
  • The fact that deep ensembles reduce false positives but still miss most anomalies suggests that uncertainty-based scoring alone is insufficient; a successful method may need to combine geometry cues, such as ground-plane inconsistency, with uncertainty.
  • A natural next step beyond the paper is a benchmark track that evaluates temporal anomaly detection, since the sequential structure is already present in STU but the reported baselines use single scans only.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This paper introduces STU, a new dataset for road anomaly segmentation in 3D LiDAR point clouds. The dataset provides 70 fully annotated sequences with anomaly instances (19 validation, 51 test), dense semantic and instance labels, synchronized eight-camera images, and 128-beam LiDAR, plus two additional 'STU-inlier' sequences for training and validation. The authors adapt several 2D anomaly segmentation methods (MaxLogit, MC Dropout, Deep Ensembles, Void Classifier, RbA) to Mask4Former-3D and report that all baselines perform poorly on the OOD task, while closed-set performance on SemanticKITTI remains reasonable. They also provide dataset statistics and supplementary analyses, including distance/size breakdowns and a 2D control experiment on the front camera. The paper claims to be the first publicly available dataset for this task with dense 3D semantic labeling, LiDAR plus camera data, and temporal sequences.

Significance. If the dataset's ground truth is trustworthy and its anomalies are genuinely out-of-distribution relative to the training data, STU would be a valuable community benchmark: it is the first public LiDAR+camera anomaly segmentation dataset with instance-level dense labels and temporal sequences, and the baseline results convincingly show that 2D-derived methods do not transfer to 3D. The paper includes honest failure reporting, per-distance AP analysis, and a 2D control experiment on the front camera, which strengthen the empirical contribution. However, the value of the benchmark is conditional on two unverified assumptions: the non-overlap between staged anomalies and training-set 'other-object'/'debris' classes, and the quality and bias of the annotation process.

major comments (2)
  1. [§4 (Training data) and Supplementary §7/§11] The OOD premise is not verified. The paper states: 'To strictly define anomalies, we analyze all other-objects present in our training set, as well as the "Movable Object.Debris" and "Pushable.Pullable" class objects in the NuScenes training set. We specifically design our dataset such that we do not have an intersection between anomalous objects and the aforementioned classes.' However, no protocol, per-category object list, or quantitative similarity check is provided. The supplementary (Section 11) shows that SemanticKITTI 'other-object' includes trash bins, garbage cans, pots, billboards, and small tables, while the staged object list (Section 7) includes indoor garbage bins, buckets, pots, bags, and similar items. Because Section 3.2 itself notes that 'current methods tend to be highly sensitive to objects that are present but ignored during training, such as those classified in the category "other object"', the reported low OOD performance could partly reflect correct inlier classification rather than genuine anomaly-detection difficulty. Please provide the per-category exclusion list, a feature-space or classifier-based check of non-overlap, and an analysis of what the trained baselines predict on the staged objects, for example confusion with the 'other-object' class.
  2. [§3.2 (Annotation process) and §4 (Metrics)] The ground-truth labels are not quantitatively validated, and the evaluation mask is partially shaped by the baselines. The annotation starts from SemanticKITTI pseudo-labels and is refined by three annotators, but no inter-annotator agreement, no error analysis, and no check of pseudo-labeler mistakes are reported. More critically, the protocol states: 'we examined the predictions of the baseline methods to see if any known objects were missed... We annotate these objects as unlabeled and ignore them in the evaluations.' Combined with the metric description ('Ignore points are removed from the scene prior to evaluation and erroneous predictions in the ignore region are not penalized'), this means that if a baseline misses an anomaly and the annotators follow this step, that anomaly is excluded from the evaluation for all methods. This creates a feedback loop between the benchmarked models and the test mask, and the direction of the bias is uncontrolled. Please quantify how many points or instances were relabeled as unlabeled through this baseline-inspection step, report inter-annotator agreement statistics, and provide a labeling protocol that does not depend on baseline predictions.
minor comments (6)
  1. [Table 1] Table 1 contains typos: '1 RBG' should be '1 RGB' and 'Augmentated' should be 'Augmented'.
  2. [References] Reference [17] spells the first author's name as 'kuefeng Du'; this appears to be a typo for 'Xuefeng Du'.
  3. [§4.1 and Supplementary §12] The paper does not report training hyperparameters (learning rate, batch size, number of epochs, validation splits) beyond the note in Supplementary §12 about lowering the learning rate; including these details would improve reproducibility.
  4. [Table 2] The OOD performance of all methods is very low (AP ≤ 5.17); the paper would benefit from per-sequence performance distributions, such as box plots over the 51 test sequences, so that readers can assess variance across object sizes and distances.
  5. [§3.2] The instruction to annotators that 'allowing for larger unlabeled regions where the annotator may be challenged' creates a potential bias toward conservative labeling; the paper should report how much of the point cloud is labeled 'unlabeled' and how this varies across sequences.
  6. [Supplementary §9.1] The 2D control experiment in Supplementary Table 5 is an important sanity check and should be considered for inclusion in the main text.

Circularity Check

0 steps flagged · score 1.0 of 10

No meaningful derivation chain to be circular; dataset and baselines are evaluated externally, with only minor self-citations that are not load-bearing.

full rationale

STU is a dataset and benchmark paper rather than a derived prediction chain. The core contribution is newly annotated LiDAR sequences, and the baseline numbers come from applying existing methods (MaxLogit, MC Dropout, Deep Ensembles, RbA, Void Classifier) to fresh test annotations, so no fitted input is renamed as a prediction. The cited Panoptic-CUDAL and Mask4Former works share authors with this paper, but they supply training data and an architectural backbone; the benchmark result is not inferred from them by construction. The only feedback loop is in label curation: Section 3.2 states that baseline predictions were used to find 'other object' points to mark as unlabeled and ignore in evaluation. That could make absolute scores slightly favorable and is a validity concern, not a circular derivation: the paper's claim is not logically equivalent to its inputs, and the evaluation mask is not itself the quantity being predicted. Score 1 reflects the minor self-citations and the labeling feedback without any self-definitional or fitted-prediction circularity.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The paper proposes no new physical entities and fits no parameters to data. The listed free parameters are manual evaluation thresholds that shape benchmark difficulty. The main unstated assumptions concern label quality and representativeness of staged anomalies; neither is quantitatively validated.

free parameters (2)
  • Minimum anomaly points for evaluation = 5 points
    Instances with fewer than five LiDAR points are excluded from point- and object-level metrics (Sec. 3.3, Fig. 6). This manual threshold removes many small and distant anomalies and directly affects benchmark difficulty.
  • Maximum evaluation range = 50 m
    All metrics are computed only for points within 50 meters, following SemanticKITTI (Sec. 3.2), even though the dataset contains labeled anomalies up to 150 m. This hand-set range excludes long-range anomalies from evaluation.
assumptions (4)
  • domain assumption Staged anomaly objects (buckets, chairs, ladders, etc.) are representative of real road debris that AVs must handle.
    Only 2 of 70 sequences are naturalistic; the rest are staged with household objects (Sec. 3.1, Supplementary Sec. 7). No evidence is given that this set covers the distribution of real-world debris.
  • domain assumption Pseudo-labeling followed by human refinement yields accurate ground-truth labels.
    The annotation process is described in Sec. 3.2 but no inter-annotator agreement, label error rate, or quality score is reported. The benchmark assumes these labels are correct.
  • domain assumption The 'unlabeled' class (objects seen in training but not supervised, e.g., parking meters, utility boxes) does not overlap with the anomaly class and can be ignored in evaluation.
    The authors selected anomaly objects to avoid intersection with SemanticKITTI 'other-object' and nuScenes debris classes (Sec. 4), but this analysis is qualitative and depends on label conventions.
  • standard math Standard metrics (PQ, UQ, AUROC, FPR@95, AP) are appropriate for evaluating road-anomaly segmentation.
    The paper adopts these metrics from prior benchmarks [4, 32, 58]; this is standard practice and not an ad hoc choice.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving." pith.science (2026). https://pith.science/paper/YB435P2I

@misc{pith2026250502148,
  author       = {Pith},
  title        = {Pith review of: Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YB435P2I}},
  note         = {Machine review of arXiv:2505.02148}
}
read the original abstract

To operate safely, autonomous vehicles (AVs) need to detect and handle unexpected objects or anomalies on the road. While significant research exists for anomaly detection and segmentation in 2D, research progress in 3D is underexplored. Existing datasets lack high-quality multimodal data that are typically found in AVs. This paper presents a novel dataset for anomaly segmentation in driving scenarios. To the best of our knowledge, it is the first publicly available dataset focused on road anomaly segmentation with dense 3D semantic labeling, incorporating both LiDAR and camera data, as well as sequential information to enable anomaly detection across various ranges. This capability is critical for the safe navigation of autonomous vehicles. We adapted and evaluated several baseline models for 3D segmentation, highlighting the challenges of 3D anomaly detection in driving environments. Our dataset and evaluation code will be openly available, facilitating the testing and performance comparison of different approaches.

Figures

Figures reproduced from arXiv: 2505.02148 by the authors.

Figure 1
Figure 1. We present Spotting the Unexpected (STU) a novel anomaly segmentation dataset for autonomous driving. The dataset contains semantic and instance labels for out-of-distribution (OOD) objects, and includes surround-view setup with synchronized cameras. Abstract To operate safely, autonomous vehicles (AVs) need to detect and handle unexpected objects or anomalies on the road. While significant research exists for anoma… view at source ↗
Figure 2
Figure 2. Data collection conducted in a naturalistic manner (a) and controlled environment (b) with objects on the road. ibration process repeated for each camera. This setup fol￾lows the configuration discussed in Panoptic-CUDAL [55]. We refer to the Supplementary Material for detailed infor￾mation on the vehicle setup. 3.1. Data collection. The data was collected using two sets of conditions: one in a naturalistic environm… view at source ↗
Figure 3
Figure 3. Different anomalies in the STU dataset. Different objects on the road used for staged data collection. We pick objects such that have no intersection with the inlier dataset and place them on roads in different locations and illumination conditions. Objects might touch each other, be very small, as large as a chair or a surf board, and could cause an accident if a car would drive over them. manticKITTI dataset, whic… view at source ↗
Figures from the paper (15 more)
Figure 4
Figure 4. Figure 4: Anomaly Instance Properties. A typical recorded anomaly has on average less then 50 points per sequence (a), with less then 300 (b) points at maximum, and a maximum height below one meter (c). We record up to nine individual anomaly instances in the same sequence (d). …
Figure 5
Figure 5. Figure 5: Distribution of anomalies along the vehicle’s reference frame. Most of the points appear around the vehicle. 1 10 100 1,000 1 101 102 Number of Points for an Anomaly Instance Distance to Anomaly Instance (meters) [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: We follow SemanticKitti’s maximum distance thresh￾old and evaluate instances that are within a 50 meter radius of the vehicle. In addition, we constrain evaluation of point and object level metrics to a minimum of 5 points. adopt several baseline methods from the 2D an…
Figure 7
Figure 7. Figure 7: Example of a failure case anomaly segmentation. For the chair on the road, labeled as anomaly in upper left, model pre￾dicts “other-vehicle” class, in lower left, with a high certainty, as indicated by MaxLogit scores on the bottom right. readers to the supplementary m…
Figure 9
Figure 9. Figure 9: shows the alignment between the LiDAR point cloud and an image captured by the front camera, illustrat￾ing the accuracy of the calibration process [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]
Figure 8
Figure 8. Figure 8: Sensor setup of the data collection vehicle. The field of view for the 60-degree and 120-degree cameras is represented in purple and blue, respectively. 6.1. Extrinsic Calibration The camera positions on the vehicle were determined through a LiDAR-camera calibration pr…
Figure 12
Figure 12. Figure 12: Patchwork++ performance in a wide environment [PITH_FULL_IMAGE:figures/full_fig_p013_12.png]
Figure 13
Figure 13. Figure 13: Patchwork++ performance in a narrow urban street. 8. Low Performance of the 3D Models 8.1. Relation of Performance to Distance and Size We calculated the AP metric for different distance thresh￾olds, as shown in [PITH_FULL_IMAGE:figures/full_fig_p013_13.png]
Figure 11
Figure 11. Figure 11: Anonymization of camera images. 7.2. Ground Plane Segmentation One of the popular approaches for anomaly detection in the point-cloud domain involves applying ground-plane re￾moval algorithms to reduce the search space. We used Patchwork++ [35] to remove the ground pl…
Figure 14
Figure 14. Figure 14: Deep Ensembles AP for differently sized objects over validation and test datasets. Method 0–10m 10–20m 20–30m 30–40m 40–50m Deep Ensemble [34] 7.63 8.49 3.42 0.38 0.03 MC Dropout [50] 0.16 0.53 0.06 0.04 0.01 Max Logit [24] 2.25 1.53 1.20 0.27 0.01 Void Classifier [4]…
Figure 15
Figure 15. Figure 15: Data Annotation Example: Each color represents a specific label — Purple for inlier, Green for anomaly, and Black for void. Boxes represent instance boundaries [PITH_FULL_IMAGE:figures/full_fig_p016_15.png]
Figure 16
Figure 16. Figure 16: Visualization of the proposed dataset with anomaly labels, instance labels, inlier class predictions, and anomaly scores of the selected anomaly methods [PITH_FULL_IMAGE:figures/full_fig_p017_16.png]
Figure 17
Figure 17. Figure 17: Example of the other-object class: a billboard, a smaller billboard, a phone booth, and a small table, all of which belong to the other-object class [PITH_FULL_IMAGE:figures/full_fig_p018_17.png]
Figure 18
Figure 18. Figure 18: Example of the other-object class: a car dealership sign and two garbage cans, all belonging to the other-object class [PITH_FULL_IMAGE:figures/full_fig_p018_18.png]
Figure 19
Figure 19. Figure 19: Example of the other-object class: a potted plant and a power adapter, all of which belong to the other-object class [PITH_FULL_IMAGE:figures/full_fig_p018_19.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

68 extracted references · 65 canonical work pages

  1. [1]

    Se- manticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences

    Jens Behley, Martin Garbade, Andres Milioto, Jan Quen- zel, Sven Behnke, Cyrill Stachniss, and Juergen Gall. Se- manticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences. In International Conference on Com- puter Vision (ICCV), 2019. 2, 3, 4, 6, 7, 8

  2. [2]

    Simultaneous Semantic Segmentation and Outlier Detection in Presence of Domain Shift

    Petra Bevandi ´c, Ivan Kreˇso, Marin Orˇsi´c, and Siniˇsa ˇSegvi´c. Simultaneous Semantic Segmentation and Outlier Detection in Presence of Domain Shift. In German Conference on Pat- tern Recognition (GCPR), 2019. 3

  3. [3]

    Nieto, Roland Y

    Hermann Blum, Paul-Edouard Sarlin, Juan I. Nieto, Roland Y . Siegwart, and C ´esar Cadena. Fishyscapes: A benchmark for safe semantic segmentation in autonomous driving. International Conference on Computer Vision Work- shop (ICCV’W), 2019. 3

  4. [4]

    The Fishyscapes Benchmark: Measuring Blind Spots in Semantic Segmentation

    Hermann Blum, Paul-Edouard Sarlin, Juan Nieto, Roland Siegwart, and Cesar Cadena. The Fishyscapes Benchmark: Measuring Blind Spots in Semantic Segmentation. Interna- tional Journal on Computer Vision (IJCV), 2021. 2, 3, 4, 6, 7

  5. [5]

    Perception datasets for anomaly detec- tion in autonomous driving: A survey

    Daniel Bogdoll, Svenja Uhlemeyer, Kamil Kowol, and J Marius Z ¨ollner. Perception datasets for anomaly detec- tion in autonomous driving: A survey. In Intelligent Vehicles Symposium (IV), 2023. 2, 3

  6. [6]

    Marius Z ¨ollner

    Daniel Bogdoll, Iramm Hamdard, Lukas Namgyu R ¨oßler, Felix Geisler, Muhammed Bayram, Felix Wang, Jan Imhof, Miguel de Campos, Anushervon Tabarov, Yitian Yang, Hanno Gottschalk, and J. Marius Z ¨ollner. AnoV ox: A Benchmark for Multimodal Anomaly Detection in Au- tonomous Driving. In ECCV 2024 W-CODA workshop ,

  7. [7]

    Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom

    Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom. nuScenes: A multi- modal dataset for autonomous driving. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2, 3

  8. [8]

    Open-set 3D Object Detection

    Jun Cen, Peng Yun, Junhao Cai, Michael Yu Wang, and Ming Liu. Open-set 3D Object Detection. In International Con- ference on 3D Vision (3DV), 2021. 3

Show all 68 references
  1. [9]

    Chakravarthy, Meghana Reddy Ganesina, Peiyun Hu, Laura Leal-Taixe, Shu Kong, Deva Ramanan, and Aljosa Osep

    Anirudh S. Chakravarthy, Meghana Reddy Ganesina, Peiyun Hu, Laura Leal-Taixe, Shu Kong, Deva Ramanan, and Aljosa Osep. Lidar Panoptic Segmentation in an Open World. In- ternational Journal on Computer Vision (IJCV), 2024. 3

  2. [10]

    SegmentMeIfYou- Can: A Benchmark for Anomaly Segmentation

    Robin Chan, Krzysztof Lis, Svenja Uhlemeyer, Hermann Blum, Sina Honari, Roland Siegwart, Pascal Fua, Math- ieu Salzmann, and Matthias Rottmann. SegmentMeIfYou- Can: A Benchmark for Anomaly Segmentation. In Proceed- ings of the Neural Information Processing Systems Track on Dat...

  3. [11]

    Entropy maximization and meta classification for out-of- distribution detection in semantic segmentation

    Robin Chan, Matthias Rottmann, and Hanno Gottschalk. Entropy maximization and meta classification for out-of- distribution detection in semantic segmentation. In Interna- tional Conference on Computer Vision (ICCV), 2021. 1, 3

  4. [12]

    Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation

    Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In European Conference on Computer Vision (ECCV), 2018. 3

  5. [13]

    Schwing, Alexan- der Kirillov, and Rohit Girdhar

    Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexan- der Kirillov, and Rohit Girdhar. Masked-attention Mask Transformer for Universal Image Segmentation. In Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,

  6. [14]

    Schwing, and Alexander Kir- illov

    Bowen Cheng, Alexander G. Schwing, and Alexander Kir- illov. Per-pixel classification is not all you need for seman- tic segmentation. In Neural Information Processing Systems (NeurIPS), 2021. 3, 6

  7. [15]

    The cityscapes dataset for semantic urban scene understanding

    Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,

  8. [16]

    Outlier detec- tion by ensembling uncertainty with negative objectness

    Anja Deli ´c, Matej Grcic, and Sini ˇsa ˇSegvi´c. Outlier detec- tion by ensembling uncertainty with negative objectness. In British Machine Vision Conference (BMVC), 2024. 1, 3

  9. [17]

    V os: Learning what you don’t know by virtual outlier synthe- sis

    kuefeng Du, Zhaoning Wang, Mu Cai, and Yixuan Li. V os: Learning what you don’t know by virtual outlier synthe- sis. In International Conference on Learning Representa- tions (ICLR), 2021. 1

  10. [18]

    Segmenting Known Objects and Unseen Unknowns without Prior Knowledge

    Stefano Gasperini, Alvaro Marcos-Ramiro, Michael Schmidt, Nassir Navab, Benjamin Busam, and Fed- erico Tombari. Segmenting Known Objects and Unseen Unknowns without Prior Knowledge. In International Conference on Computer Vision (ICCV), 2023. 1, 3, 5

  11. [19]

    Vision meets robotics: The KITTI dataset

    A Geiger, P Lenz, C Stiller, and R Urtasun. Vision meets robotics: The KITTI dataset. In The International Journal of Robotics Research, 2013. 3

  12. [20]

    Dense open- set recognition with synthetic outliers generated by real nvp

    Matej Grcic, Petra Bevandic, and Sinisa Segvic. Dense open- set recognition with synthetic outliers generated by real nvp. ArXiv, abs/2011.11094, 2020. 3

  13. [21]

    Densehy- brid: Hybrid anomaly detection for dense open-set recogni- tion

    Matej Grci ´c, Petra Bevandi ´c, and Sini ˇsa ˇSegvi´c. Densehy- brid: Hybrid anomaly detection for dense open-set recogni- tion. In European Conference on Computer Vision (ECCV),

  14. [22]

    Madhava Krishna

    Krishnam Gupta, Syed Ashar Javed, Vineet Gandhi, and K. Madhava Krishna. MergeNet: A Deep Net Architecture for Small Obstacle Discovery. In International Conference on Robotics and Automation (ICRA), 2018. 2

  15. [23]

    SODA10M: A Large- Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving

    Jianhua Han, Xiwen Liang, Hang Xu, Kai Chen, Lanqing Hong, Jiageng Mao, Chaoqiang Ye, Wei Zhang, Zhenguo Li, Xiaodan Liang, and Chunjing Xu. SODA10M: A Large- Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving. In Proceedings of the Neural Informa- t...

  16. [24]

    A Baseline for Detect- ing Misclassified and Out-of-Distribution Examples in Neu- ral Networks

    Dan Hendrycks and Kevin Gimpel. A Baseline for Detect- ing Misclassified and Out-of-Distribution Examples in Neu- ral Networks. In International Conference on Learning Rep- resentations (ICLR), 2018. 2, 3, 6, 7, 4

  17. [25]

    Scaling Out-of-Distribution Detection for Real- World Settings

    Dan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou, Joe Kwon, Mohammadreza Mostajabi, Jacob Steinhardt, and Dawn Song. Scaling Out-of-Distribution Detection for Real- World Settings. In International Conference on Machine Learning (ICML), 2022. 2, 3

  18. [26]

    Scaling out-of-distribution detection for real- world settings

    Dan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou, Mohammadreza Mostajabi, Jacob Steinhardt, and Dawn Xi- aodong Song. Scaling out-of-distribution detection for real- world settings. In International Conference on Machine Learning, 2022. 8

  19. [27]

    Generalized ODIN: Detecting Out-of-Distribution Image Without Learning From Out-of-Distribution Data

    Yen-Chang Hsu, Yilin Shen, Hongxia Jin, and Zsolt Kira. Generalized ODIN: Detecting Out-of-Distribution Image Without Learning From Out-of-Distribution Data . In Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 8

  20. [28]

    Czarnecki

    Chengjie Huang, Van Duong Nguyen, Vahdat Abdelzad, Christopher Gus Mannes, Luke Rowe, Benjamin Therien, Rick Salay, and K. Czarnecki. Out-of-distribution detection for lidar-based 3d object detection. IEEE Intelligent Trans- portation Systems Conference (ITSC), 2022. 3

  21. [29]

    Deepprivacy2: To- wards realistic full-body anonymization

    H ˚akon Hukkel ˚as and Frank Lindseth. Deepprivacy2: To- wards realistic full-body anonymization. In Winter Confer- ence on Applications of Computer Vision (WACV), 2023. 4, 2

  22. [30]

    Exemplar-Based Open-Set Panoptic Segmenta- tion Network

    Jaedong Hwang, Seoung Wug Oh, Joon-Young Lee, and Bo- hyung Han. Exemplar-Based Open-Set Panoptic Segmenta- tion Network. In Conference on Computer Vision and Pat- tern Recognition (CVPR), 2021. 3

  23. [31]

    Standardized Max Logits: A Simple yet Effective Approach for Identifying Unexpected Road Obsta- cles in Urban-Scene Segmentation

    Sanghun Jung, Jungsoo Lee, Daehoon Gwak, Sungha Choi, and Jaegul Choo. Standardized Max Logits: A Simple yet Effective Approach for Identifying Unexpected Road Obsta- cles in Urban-Scene Segmentation. In International Confer- ence on Computer Vision (ICCV), 2021. 1

  24. [32]

    Panoptic Segmentation

    Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Doll´ar. Panoptic Segmentation. In Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,

  25. [33]

    Revisiting Out-of-Distribution Detection in LiDAR-based 3D Object Detection

    Michael K ¨osel, Marcel Schreiber, Michael Ulrich, Claudius Gl¨aser, and Klaus Dietmayer. Revisiting Out-of-Distribution Detection in LiDAR-based 3D Object Detection. In Intelli- gent Vehicles Symposium (IV), 2024. 3

  26. [34]

    Simple and Scalable Predictive Uncertainty Es- timation using Deep Ensembles

    Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and Scalable Predictive Uncertainty Es- timation using Deep Ensembles. In Neural Information Pro- cessing Systems (NeurIPS), 2017. 2, 3, 6, 7, 4

  27. [35]

    Patch- work++: Fast and robust ground segmentation solving par- tial under-segmentation using 3D point cloud

    Seungjae Lee, Hyungtae Lim, and Hyun Myung. Patch- work++: Fast and robust ground segmentation solving par- tial under-segmentation using 3D point cloud. In Proc. IEEE/RSJ Int. Conf. Intell. Robots Syst., 2022. 3, 2

  28. [36]

    Coda: A real-world road corner case dataset for object detection in autonomous driving

    Kaican Li, Kai Chen, Haoyu Wang, Lanqing Hong, Chao- qiang Ye, Jianhua Han, Yukuai Chen, Wei Zhang, Chunjing Xu, Dit-Yan Yeung, et al. Coda: A real-world road corner case dataset for object detection in autonomous driving. In European Conference on Computer Vision (ECCV), 2022. 2, 3

  29. [37]

    GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models

    Chen Liang, Wenguan Wang, Jiaxu Miao, and Yi Yang. GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models. In Neural Information Processing Systems (NeurIPS), 2022. 1, 3

  30. [38]

    Two Video Data Sets for Tracking and Retrieval of Out of Distribution Objects

    Kira Maag, Robin Chan, Svenja Uhlemeyer, Kamil Kowol, and Hanno Gottschalk. Two Video Data Sets for Tracking and Retrieval of Out of Distribution Objects. In Asian Con- ference on Computer Vision (ACCV), 2022. 2, 3

  31. [39]

    One Million Scenes for Autonomous Driving: ONCE Dataset

    Jiageng Mao, Minzhe Niu, Chenhan Jiang, Hanxue Liang, Jingheng Chen, Xiaodan Liang, Yamin Li, Chaoqiang Ye, Wei Zhang, Zhenguo Li, Jie Yu, Hang Xu, and Chun- jing Xu. One Million Scenes for Autonomous Driving: ONCE Dataset. In Neural Information Processing Systems (NeurIPS), 2021. 3

  32. [40]

    Mask-Based Panoptic LiDAR Segmentation for Autonomous Driving

    Rodrigo Marcuzzi, Lucas Nunes, Louis Wiesmann, Jens Behley, and Cyrill Stachniss. Mask-Based Panoptic LiDAR Segmentation for Autonomous Driving. In IEEE Robotics And Automation Letters (RAL), 2023. 7, 4

  33. [41]

    Henriques, and Fatma G¨uney

    Nazir Nayal, Mısra Yavuz, Jo ˜ao F. Henriques, and Fatma G¨uney. RbA: Segmenting Unknown Regions Rejected by All. In International Conference on Computer Vision (ICCV), 2023. 1, 2, 3, 6, 7, 4

  34. [42]

    OoDIS: Anomaly Instance Segmentation Benchmark

    Alexey Nekrasov, Rui Zhou, Miriam Ackermann, Alexander Hermans, and Matthias Rottmann Bastian Leibe. OoDIS: Anomaly Instance Segmentation Benchmark. arXiv preprint arXiv:2406.11835, 2024. 2

  35. [43]

    The mapillary vistas dataset for semantic understanding of street scenes

    Gerhard Neuhold, Tobias Ollmann, Samuel Rota Bulo, and Peter Kontschieder. The mapillary vistas dataset for semantic understanding of street scenes. In International Conference on Computer Vision (ICCV), 2017. 3

  36. [44]

    Unsupervised Class-Agnostic Instance Segmentation of 3D LiDAR Data for Autonomous Vehicles

    Lucas Nunes, Xieyuanli Chen, Rodrigo Marcuzzi, Aljosa Osep, Laura Leal-Taix ´e, Cyrill Stachniss, and Jens Behley. Unsupervised Class-Agnostic Instance Segmentation of 3D LiDAR Data for Autonomous Vehicles. IEEE Robotics And Automation Letters (RAL), 2022. 3

  37. [45]

    Lost and Found: Detecting Small Road Hazards for Self-Driving Vehicles

    Peter Pinggera, Sebastian Ramos, Stefan Gehrig, Uwe Franke, Carsten Rother, and Rudolf Mester. Lost and Found: Detecting Small Road Hazards for Self-Driving Vehicles. In International Conference on Intelligent Robots and Systems (IROS), 2016. 2, 3, 4

  38. [46]

    LS-VOS: Identifying Outliers in 3D Object Detections Using Latent Space Virtual Outlier Synthesis

    Aldi Piroli, Vinzenz Dallabetta, Johannes Kopp, Marc Wa- lessa, Daniel Meissner, and Klaus Dietmayer. LS-VOS: Identifying Outliers in 3D Object Detections Using Latent Space Virtual Outlier Synthesis. In IEEE Intelligent Trans- portation Systems Conference (ITSC), 2023. 2, 3

  39. [47]

    Unmasking Anomalies in Road-Scene Segmentation

    Shyam Nandan Rai, Fabio Cermelli, Dario Fontanel, Carlo Masone, and Barbara Caputo. Unmasking Anomalies in Road-Scene Segmentation. In International Conference on Computer Vision (ICCV), 2023. 3

  40. [48]

    SeMoLi: What Moves Together Belongs Together

    Jenny Seidenschwarz, Aljo ˇsa O ˇsep, Francesco Ferroni, Si- mon Lucey, and Laura Leal-Taix ´e. SeMoLi: What Moves Together Belongs Together. In Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2024. 3

  41. [49]

    Madhava Krishna

    Aasheesh Singh, Aditya Kamireddypalli, Vineet Gandhi, and K. Madhava Krishna. LiDAR guided Small obstacle Seg- mentation. In International Conference on Intelligent Robots and Systems (IROS), 2020. 2, 3

  42. [50]

    Dropout: a simple way to prevent neural networks from overfitting

    Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. In Neu- ral Information Processing Systems (NeurIPS) , 2014. 2, 6, 7, 3, 4

  43. [51]

    Scalability in perception for autonomous driving: Waymo open dataset

    Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. InConference on Computer Vision and Pattern Recognition ...

  44. [52]

    Dashcamcleaner: Censor identifiable information in videos from dashcam recordings

    DashcamCleaner Development Team. Dashcamcleaner: Censor identifiable information in videos from dashcam recordings. https : / / github . com / tfaehse / DashcamCleaner, 2024. 4, 2

  45. [53]

    The Prevalence of Motor Vehicle Crashes Involving Road Debris, United States, 2011-2014

    Tefft, Brian C. . The Prevalence of Motor Vehicle Crashes Involving Road Debris, United States, 2011-2014. https: //aaafoundation.org/wp- content/uploads/ 2017/12/RoadDebris_REPORT_2015.pdf , 2015. [Online]. 1

  46. [54]

    Pixel-wise Energy-biased Abstention Learning for Anomaly Segmentation on Com- plex Urban Driving Scenes

    Yu Tian, Yuyuan Liu, Guansong Pang, Fengbei Liu, Yuan- hong Chen, and Gustavo Carneiro. Pixel-wise Energy-biased Abstention Learning for Anomaly Segmentation on Com- plex Urban Driving Scenes. In European Conference on Computer Vision (ECCV), 2022. 1, 3

  47. [55]

    Panoptic-CUDAL Technical Report: Ru- ral Australia Point Cloud Dataset in Rainy Conditions.arXiv preprint arXiv:2503.16378, 2025

    Tzu-Yun Tseng, Alexey Nekrasov, Malcolm Burdorf, Bas- tian Leibe, Julie Stephany Berrio Perez, Mao Shan, and Stewart Worrall. Panoptic-CUDAL Technical Report: Ru- ral Australia Point Cloud Dataset in Rainy Conditions.arXiv preprint arXiv:2503.16378, 2025. 3, 4, 6, 8

  48. [56]

    Verma, J

    S. Verma, J. S. Berrio, S. Worrall, and E. Nebot. Automatic extrinsic calibration between a camera and a 3d lidar using 3d point and plane correspondences. In IEEE Intelligent Transportation Systems Conference (ITSC), 2019. 3

  49. [57]

    KISS-ICP: In Defense of Point-to-Point ICP – Simple, Accurate, and Ro- bust Registration If Done the Right Way

    Ignacio Vizzo, Tiziano Guadagnino, Benedikt Mersch, Louis Wiesmann, Jens Behley, and Cyrill Stachniss. KISS-ICP: In Defense of Point-to-Point ICP – Simple, Accurate, and Ro- bust Registration If Done the Right Way. In IEEE Robotics And Automation Letters (RAL), 2023. 4, 2

  50. [58]

    Identifying Unknown Instances for Autonomous Driving

    Kelvin Wong, Shenlong Wang, Mengye Ren, Ming Liang, and Raquel Urtasun. Identifying Unknown Instances for Autonomous Driving. In Conference on Robot Learning (CoRL), 2019. 2, 3, 7, 8

  51. [59]

    Raod: A benchmark for road abandoned object detection from video surveillance

    Yajun Xu, Huan Hu, Xiaoya Zhu, Yibing Nan, Kai Wang, ZhaoXiang Liu, and Shiguo Lian. Raod: A benchmark for road abandoned object detection from video surveillance. IEEE Access, 2024. 2

  52. [60]

    Mask4Former: Mask Transformer for 4D Panoptic Segmentation

    Kadir Yilmaz, Jonas Schult, Alexey Nekrasov, and Bastian Leibe. Mask4Former: Mask Transformer for 4D Panoptic Segmentation. In International Conference on Robotics and Automation (ICRA), 2024. 3, 6, 7, 4

  53. [61]

    Bdd100k: A diverse driving dataset for heterogeneous multitask learning

    Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Dar- rell. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 3 Spotting the U...

  54. [62]

    Hardware Setup The sensors and hardware included in the data collection platform are as follows: • 5 SF3325 automotive GMSL cameras (ONSEMI CMOS image sensor AR0231), SEKONIX ultra high-resolution lens with 60 horizontal and 38 vertical FOV , images cap- tured at a resolution ...

  55. [63]

    The point cloud captures the three- dimensional structures of objects at varying distances

    Data Collection For staged data collection, we used a diverse collection of objects, including buckets, indoor garbage bins, brooms, chairs, pots, stuffed animals, balloons, balls, backpacks, bags, pillows, shoes, umbrellas, hats, yoga mats, helmets, swimming noodles, tissue b...

  56. [64]

    Relation of Performance to Distance and Size We calculated the AP metric for different distance thresh- olds, as shown in Table 4

    Low Performance of the 3D Models 8.1. Relation of Performance to Distance and Size We calculated the AP metric for different distance thresh- olds, as shown in Table 4. In the lower ranges, from 0 to 10, and from 10 to 20 range models perform better, then at other distances. N...

  57. [65]

    unlabeled

    Results on Validation Datasets We show results for the SemanticKITTI [1] validation set in Table 6 and our dataset in Table 7. For the OOD validation set, we evaluate in three sequences and provide scores in Table 8. Method Aux Data AUROC ↑ FPR@95 ↓ AP↑ DenseHybrid [21] ✗ 87.0...

  58. [66]

    Anomaly points are cyan, unlabeled regions are black, and inliers are pur- ple

    Annotation and Qualitative Examples We visualize the annotation interface with an example of a correctly annotated scene in the figure 15. Anomaly points are cyan, unlabeled regions are black, and inliers are pur- ple. We provide further visualizations of the dataset and the p...

  59. [67]

    The other-object class consists of many mis- cellaneous items, including trash bins, advertisement posts, and small pots

    SemanticKITTI Other-object Examples Several examples of the other-object class in the Se- manticKITTI dataset can be seen in Figure 17, Figure 18, and Figure 19. The other-object class consists of many mis- cellaneous items, including trash bins, advertisement posts, and small...

  60. [68]

    This also occurred during training runs solely on Panoptic-CUDAL

    Note on Training Initially, jointly training with both SemanticKITTI and Panoptic-CUDAL led to diverging losses for Mask4Former- 3D. This also occurred during training runs solely on Panoptic-CUDAL. Lowering the preset learning rate from 0.0004 to 0.0002 was enough to mitigate...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.