REVIEW 4 major objections 6 minor 93 references
BelHouse3D: A Benchmark Dataset for Assessing Occlusion Robustness in 3D Point Cloud Semantic Segmentation
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read BelHouse3D, a synthetic indoor point cloud dataset built from 32 Belgian houses, shows that occlusion drops fully supervised semantic segmentation mIoU by 30-49% while few-shot methods remain robust.
desk verdict A useful dataset idea with a clever occlusion construction, but the load-bearing claim of real-world alignment is unvalidated and the dataset is not released. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the pipeline that turns real scans into paired clean and occluded synthetic point clouds. RGB-D frames captured with a handheld sensor are reconstructed into real point clouds; those reconstructions guide the layout of synthetic rooms and the placement of predefined object models in a 3D modeling program. To create the OOD test set, the synthetic scene is filtered through the original recording viewpoints, keeping only the points that would have been visible from those camera positions. This single filtering step is what converts an approximately arranged synthetic scene into occlusion patterns that mimic real-world visibility, and it is also the step that the paper's conclusions depend on.
What would settle it
Compare the occluded regions in BelHouse3D's OOD test set with the actual missing-point patterns in the real reconstructed scans of the same houses when viewed from the same viewpoints; if the shapes and frequencies of partial object occlusion do not match, the simulated OOD shift is not faithful. A more direct test is to train a model on the clean BelHouse3D set and evaluate it on real occluded scans, checking whether the mIoU drop falls in the paper's reported 30-49 percent range.
Extended reading notes
Core claim
The paper claims that occlusion, a common and unavoidable condition of real-world indoor point clouds, constitutes a significant out-of-distribution shift that current fully supervised segmentation models are not robust to. On BelHouse3D, PointNet++ drops from 71.97 to 36.56 mIoU (-49%), DGCNN from 72.57 to 40.34 (-44%), Stratified Transformer from 79.12 to 44.54 (-44%), and Point TransformerV2 from 82.32 to 55.95 (-32%), with PointNet dropping from 38.38 to 26.97 (-30%). Since overall accuracy declines far less than mIoU, the paper argues occlusion disproportionately hurts smaller object classes rather than large building structures. The few-shot experiments show prototype-based and attention-based few-shot models are markedly more stable, with MPTI and AttMPTI improving on novel classes under occlusion in several configurations. The authors position BelHouse3D as the first dedicated benchmark for occlusion-based OOD generalization in indoor 3D point cloud segmentation.
Load-bearing premise
The benchmark's conclusions assume that filtering the synthetic scenes through the original recording viewpoints reproduces the statistical structure of real occlusions, even though object placement in the synthetic scenes is only approximate and does not strictly adhere to exact positions.
Editorial extensions
If this is right
- Fully supervised point-based segmentation models should be evaluated under occlusion-style OOD shifts, because IID test scores overstate real-world reliability.
- The reported 30-49 percent mIoU drops provide a quantitative baseline that any proposed occlusion-robust method should beat on this benchmark.
- Few-shot learning with prototype or attention mechanisms appears to be a more stable route to occlusion robustness than full supervision.
- The gap between mIoU and overall accuracy under occlusion indicates that robustness efforts should concentrate on small object classes, not building structures.
Reading between the lines
- A natural extension the authors do not run is to apply the same viewpoint-filtering recipe to real indoor datasets, which would test whether synthetic occlusion statistics match real ones without relying on approximate object placement.
- If the few-shot robustness finding generalizes, it suggests occlusion resistance may come from transferable geometric prototypes rather than from large labeled training sets.
- Because BelHouse3D currently stores only XYZ coordinates, the benchmark isolates geometry-based robustness; adding color or surface normals could change the measured degradation magnitudes.
- The dataset could also support sim-to-real transfer studies by pretraining on clean synthetic scenes and fine-tuning on real occluded scans, a regime the paper does not report.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces BelHouse3D, a synthetic point cloud dataset for indoor semantic segmentation, built from real-world references of 32 Belgian houses. It contains clean, fully labeled point clouds for training and in-distribution testing, plus an out-of-distribution (OOD) test set created by filtering synthetic points through the original RealSense viewpoints to simulate occlusion. The authors benchmark five fully supervised point-based segmentation methods (PointNet, PointNet++, DGCNN, Stratified Transformer, Point Transformer V2) and four few-shot segmentation methods (ProtoNet, AttProto, MPTI, AttMPTI) on both IID and OOD test sets, reporting mIoU and OA. Their main empirical finding is that fully supervised models degrade substantially under occlusion (mIoU drops of 32–49% in Table 2), while few-shot methods appear comparatively more robust.
Significance. If the dataset is publicly released and the OOD construction is quantitatively validated, BelHouse3D would fill a clear gap: a dedicated 3D indoor point cloud benchmark for occlusion robustness. The paper provides initial baselines and a concrete evaluation protocol, which could be reused by the community. The finding that fully supervised methods lose large amounts of mIoU under occlusion, while few-shot methods show smaller or even positive changes, is interesting and falsifiable. However, the current manuscript does not release the dataset or code, and the core assumption that the viewpoint-filtered synthetic occlusions faithfully reproduce real occlusions is not verified, thus limiting the immediate significance and reliability of the reported numbers.
major comments (4)
- [Sec. 3.2 and Sec. 3.3] The OOD test set is the load-bearing contribution: every downstream conclusion about model degradation and few-shot robustness depends on its realism. The construction filters synthetic points using the original RealSense viewpoints (Sec. 3.3), but Sec. 3.2 states that object placement 'does not strictly adhere to exact object positioning.' The paper provides no quantitative evidence that the synthetic scene geometry aligns with the real reconstructions well enough for these visibility masks to be meaningful. Without an alignment error metric or a comparison of occlusion statistics (e.g., per-class visibility rates, occlusion pattern distributions) between the synthetic OOD set and the real scans, the claim that the occlusions 'closely mimic real-world scenarios' is unsubstantiated. Please add such validation or explicitly reframe the OOD set as a synthetic, controlled occlusion model whose realism is not yet established.
- [Throughout (data availability)] This is a benchmark dataset paper, but no URL, download link, or data availability statement is provided for BelHouse3D, and no code release is mentioned. Without access to the dataset, readers cannot reproduce the tables, verify the OOD construction, or use the benchmark as a community resource. The manuscript should include a clear statement on where and under what terms the dataset and evaluation code will be released.
- [Tables 2 and 3] The paper reports a single run per method for each setting, with no error bars, standard deviations, or statistical tests. The central comparative claim in Sec. 4.2 that 'few-shot learning methods exhibit better robustness' rests on differences that are often small (e.g., Table 3, 1-way 1-shot: AttMPTI gains +8.28% on Novel, but ProtoNet gains only -0.09% on Novel and several other changes are within 1-2%). Without repeated runs or a measure of variance, these differences may not be significant. Please report mean and standard deviation over at least three seeds, or otherwise demonstrate that the observed robustness gap is not due to noise.
- [Sec. 4.2 (FSL setup)] The description of the few-shot support/query construction is ambiguous. The text says 'N × 5 samples are selected from the entire sample list for each class and designated as the support set, with the remaining samples forming the query set,' but it is unclear what a 'sample' is (a point cloud block, a subcloud, or a fixed-size point set) and whether support and query points are drawn from the same point clouds or different ones. This ambiguity affects the reproducibility of Table 3 and should be clarified with precise definitions.
minor comments (6)
- [Table 1] ScanNet200 is listed as synthetic (S) in the R/S column, but it is a real-world dataset. This mislabeling should be corrected to accurately situate BelHouse3D relative to prior work.
- [Sec. 2 (Related Work)] The text refers to 'SceneNet [30]' but reference [30] is SceneNN (Hua et al.). SceneNet is a different dataset (McCormac et al.); please correct the citation or the dataset name.
- [Fig. 2 caption] The caption mentions 'bubble sizes represent the number of points in each class,' but the figure appears to be a bar graph. Please align the caption with the actual visualization or clarify what the bubbles refer to.
- [Sec. 3.2] The process of 'sampling' points from Blender surfaces is not described in enough detail: no density, number of points per scene, or sampling strategy is specified. These details are important for reproducibility and for understanding the IID/OOD point distributions.
- [Sec. 3.3] The paper states that occluded test data are generated 'using the original viewpoints' but does not explain how visibility is computed (e.g., ray casting, depth buffers, or projection). Adding this implementation detail would strengthen reproducibility.
- [Sec. 5 (Limitations)] The limitations section mentions the future expansion of classes and attributes but does not acknowledge the approximate object placement or the lack of validation of the OOD realism; consider adding these as limitations.
Circularity Check
No significant circularity found; the dataset and benchmark results are empirical constructions and measurements, not derivations from their own assumptions.
full rationale
This is a dataset and empirical benchmark paper. The load-bearing claims are that BelHouse3D provides a synthetic point cloud dataset with an occlusion-based OOD test set, and that fully supervised models degrade substantially while few-shot methods are more robust under that OOD setting. These claims are supported by dataset construction and experimental measurements, not by a derivation that assumes its own conclusion. The OOD test set is produced by filtering synthetic points through the original RealSense viewpoints (Section 3.3), a physical simulation step that does not use the evaluated models' predictions or any fitted parameter. The benchmark results in Tables 2 and 3 are measured outcomes on held-out houses and classes, so they are externally checkable and falsifiable. The only self-referential aspect is that the authors constructed the test data themselves, which is standard and not circular. The paper's admission that object referencing 'does not strictly adhere to exact object positioning' is a validity concern about whether the occlusion simulation faithfully reproduces real-world occlusions, but it is not an instance of a prediction being equivalent to an input by construction. No self-citation chain is load-bearing, no known result is merely renamed, and no fitted input is relabeled as a prediction. Accordingly, the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Synthetic point clouds sampled from Blender surfaces are a valid proxy for real-world indoor point clouds for training and evaluating segmentation models.
- ad hoc to paper Occlusion introduced by filtering synthetic points through original RealSense viewpoints approximates the occlusion distribution of real point clouds.
- domain assumption Real-world references from 32 Belgian houses make the dataset fair and representative.
- domain assumption Performance drop under synthetic occlusion predicts real-world OOD degradation.
Cite this review
Pith. "Pith review of BelHouse3D: A Benchmark Dataset for Assessing Occlusion Robustness in 3D Point Cloud Semantic Segmentation." pith.science (2026). https://pith.science/paper/NQXLRTG4
@misc{pith2026241113251,
author = {Pith},
title = {Pith review of: BelHouse3D: A Benchmark Dataset for Assessing Occlusion Robustness in 3D Point Cloud Semantic Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/NQXLRTG4}},
note = {Machine review of arXiv:2411.13251}
}
read the original abstract
Large-scale 2D datasets have been instrumental in advancing machine learning; however, progress in 3D vision tasks has been relatively slow. This disparity is largely due to the limited availability of 3D benchmarking datasets. In particular, creating real-world point cloud datasets for indoor scene semantic segmentation presents considerable challenges, including data collection within confined spaces and the costly, often inaccurate process of per-point labeling to generate ground truths. While synthetic datasets address some of these challenges, they often fail to replicate real-world conditions, particularly the occlusions that occur in point clouds collected from real environments. Existing 3D benchmarking datasets typically evaluate deep learning models under the assumption that training and test data are independently and identically distributed (IID), which affects the models' usability for real-world point cloud segmentation. To address these challenges, we introduce the BelHouse3D dataset, a new synthetic point cloud dataset designed for 3D indoor scene semantic segmentation. This dataset is constructed using real-world references from 32 houses in Belgium, ensuring that the synthetic data closely aligns with real-world conditions. Additionally, we include a test set with data occlusion to simulate out-of-distribution (OOD) scenarios, reflecting the occlusions commonly encountered in real-world point clouds. We evaluate popular point-based semantic segmentation methods using our OOD setting and present a benchmark. We believe that BelHouse3D and its OOD setting will advance research in 3D point cloud semantic segmentation for indoor scenes, providing valuable insights for the development of more generalizable models.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops
Agustsson, E., Timofte, R.: Ntire 2017 challenge on single image super-resolution: Dataset and study. In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops. pp. 126–135 (2017)
2017
-
[2]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Armeni, I., Sener, O., Zamir, A.R., Jiang, H., Brilakis, I., Fischer, M., Savarese, S.: 3d semantic parsing of large-scale indoor spaces. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1534–1543 (2016)
2016
-
[3]
In: Proceedings of the IEEE/CVF international conference on computer vision
Behley, J., Garbade, M., Milioto, A., Quenzel, J., Behnke, S., Stachniss, C., Gall, J.: Semantickitti: A dataset for semantic scene understanding of lidar sequences. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 9297–9307 (2019)
2019
-
[4]
In: International Conference on Machine Learning
Bitterwolf, J., Müller, M., Hein, M.: In or out? fixing imagenet out-of-distribution detection evaluation. In: International Conference on Machine Learning. pp. 2471–
-
[5]
blender.org/, accessed: 10-08-2024
Blender: free and open-source 3d computer graphics software tool.https://www. blender.org/, accessed: 10-08-2024
2024
-
[6]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O.: nuscenes: A multimodal dataset for autonomous driving. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 11621–11631 (2020)
2020
-
[7]
arXiv preprint arXiv:1709.06158 (2017)
Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., Zhang, Y.: Matterport3d: Learning from rgb-d data in indoor envi- ronments. arXiv preprint arXiv:1709.06158 (2017)
arXiv 2017
-
[8]
arXiv preprint arXiv:1512.03012 (2015)
Chang, A.X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., Su, H., et al.: Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012 (2015)
arXiv 2015
Show all 93 references
-
[9]
arXiv preprint arXiv:2203.09065 (2022)
Chen, M., Hu, Q., Yu, Z., Thomas, H., Feng, A., Hou, Y., McCullough, K., Ren, F., Soibelman, L.: Stpls3d: A large-scale synthetic and real aerial photogrammetry 3d point cloud dataset. arXiv preprint arXiv:2203.09065 (2022)
2022
-
[10]
In: 2019 International Conference on 3D Vision (3DV)
Chiang, H.Y., Lin, Y.L., Liu, Y.C., Hsu, W.H.: A unified point-based framework for 3d segmentation. In: 2019 International Conference on 3D Vision (3DV). pp. 155–163. IEEE (2019) BelHouse3D 15
2019
-
[11]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Cho, J., Li, L., Yang, Z., Gan, Z., Wang, L., Bansal, M.: Diagnostic benchmark and iterative inpainting for layout-guided image generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 5280– 5289 (2024)
2024
-
[12]
In: 2017 international joint conference on neural networks (IJCNN)
Cohen, G., Afshar, S., Tapson, J., Van Schaik, A.: Emnist: Extending mnist to handwritten letters. In: 2017 international joint conference on neural networks (IJCNN). pp. 2921–2926. IEEE (2017)
2017
-
[13]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Dai,A.,Chang,A.X.,Savva,M.,Halber,M.,Funkhouser,T.,Nießner,M.:Scannet: Richly-annotated 3d reconstructions of indoor scenes. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 5828–5839 (2017)
2017
-
[14]
In: Proceedings of the European Conference on Computer Vision (ECCV)
Dai, A., Nießner, M.: 3dmv: Joint 3d-multi-view prediction for 3d semantic scene segmentation. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 452–468 (2018)
2018
-
[15]
NeurIPS Datasets and Benchmarks 2(6), 16 (2021)
Dehghan, A., Baruch, G., Chen, Z., Feigin, Y., Fu, P., Gebauer, T., Kurz, D., Dimry, T., Joffe, B., Schwartz, A., et al.: Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. NeurIPS Datasets and Benchmarks 2(6), 16 (2021)
2021
-
[16]
In: 2009 IEEE conference on computer vision and pattern recognition
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large- scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition. pp. 248–255. Ieee (2009)
2009
-
[17]
arXiv preprint arXiv:2405.17419 (2024)
Dong, H., Zhao, Y., Chatzi, E., Fink, O.: Multiood: Scaling out-of-distribution detection for multiple modalities. arXiv preprint arXiv:2405.17419 (2024)
2024 arXiv
-
[18]
https://www.dotproduct3d.com/ dot3d.html, accessed: 10-08-2024
DotProduct: The dot3d scanning platform. https://www.dotproduct3d.com/ dot3d.html, accessed: 10-08-2024
2024
-
[19]
International journal of computer vision88, 303–338 (2010)
Everingham,M.,VanGool,L.,Williams,C.K.,Winn,J.,Zisserman,A.:Thepascal visual object classes (voc) challenge. International journal of computer vision88, 303–338 (2010)
2010
-
[20]
In: 2004 conference on computer vision and pattern recognition workshop
Fei-Fei, L., Fergus, R., Perona, P.: Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object cate- gories. In: 2004 conference on computer vision and pattern recognition workshop. pp. 178–178. IEEE (2004)
2004
-
[21]
arXiv preprint arXiv:2302.11893 (2023)
Galil, I., Dabbah, M., El-Yaniv, R.: A framework for benchmarking class- out-of-distribution detection and its application to imagenet. arXiv preprint arXiv:2302.11893 (2023)
2023 arXiv
-
[22]
In: 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
Garcia-Garcia, A., Martinez-Gonzalez, P., Oprea, S., Castro-Vargas, J.A., Orts- Escolano, S., Garcia-Rodriguez, J., Jover-Alvarez, A.: The robotrix: An extremely photorealistic and very-large-scale indoor dataset of sequences with robot trajec- tories and interactions. In: 201...
2018
-
[23]
Communications of the ACM 63(11), 139–144 (2020)
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial networks. Communications of the ACM 63(11), 139–144 (2020)
2020
-
[24]
Advances in neural information processing systems 30 (2017)
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., Courville, A.C.: Improved training of wasserstein gans. Advances in neural information processing systems 30 (2017)
2017
-
[25]
IEEE transactions on pattern analysis and machine intelligence 43(12), 4338–4364 (2020)
Guo, Y., Wang, H., Hu, Q., Liu, H., Liu, L., Bennamoun, M.: Deep learning for 3d point clouds: A survey. IEEE transactions on pattern analysis and machine intelligence 43(12), 4338–4364 (2020)
2020
-
[26]
net: A new large-scale point cloud classification benchmark
Hackel, T., Savinov, N., Ladicky, L., Wegner, J.D., Schindler, K., Pollefeys, M.: Semantic3d. net: A new large-scale point cloud classification benchmark. arXiv preprint arXiv:1704.03847 (2017) 16 U. Raman Kumar et al
2017 arXiv
-
[27]
arXiv preprint arXiv:1903.12261 (2019)
Hendrycks,D.,Dietterich,T.:Benchmarkingneuralnetworkrobustnesstocommon corruptions and perturbations. arXiv preprint arXiv:1903.12261 (2019)
2019 arXiv
-
[28]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Hu, Q., Yang, B., Khalid, S., Xiao, W., Trigoni, N., Markham, A.: Towards se- mantic segmentation of urban-scale 3d point clouds: A dataset, benchmarks and challenges. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 4977–4987 (2021)
2021
-
[29]
In: Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition
Hu, Q., Yang, B., Xie, L., Rosa, S., Guo, Y., Wang, Z., Trigoni, N., Markham, A.: Randla-net: Efficient semantic segmentation of large-scale point clouds. In: Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 11108–11117 (2020)
2020
-
[30]
In: 2016 fourth international conference on 3D vision (3DV)
Hua, B.S., Pham, Q.H., Nguyen, D.T., Tran, M.K., Yu, L.F., Yeung, S.K.: Scenenn: A scene meshes dataset with annotations. In: 2016 fourth international conference on 3D vision (3DV). pp. 92–101. Ieee (2016)
2016
-
[31]
Jagadeesh, A.V., Gardner, J.L.: Texture-like representation of objects in human vi- sualcortex.ProceedingsoftheNationalAcademyofSciences 119(17),e2115302119 (2022)
2022
-
[32]
In: Pro- ceedings of the IEEE/CVF international conference on computer vision workshops
Jaritz, M., Gu, J., Su, H.: Multi-view pointnet for 3d scene understanding. In: Pro- ceedings of the IEEE/CVF international conference on computer vision workshops. pp. 0–0 (2019)
2019
-
[33]
International Journal of Computer Vision126, 920–941 (2018)
Jiang, C., Qi, S., Zhu, Y., Huang, S., Lin, J., Yu, L.F., Terzopoulos, D., Zhu, S.C.: Configurable 3d scene synthesis and 2d image rendering with per-pixel ground truth using stochastic grammars. International Journal of Computer Vision126, 920–941 (2018)
2018
-
[34]
arXiv preprint arXiv:1807.00652 (2018)
Jiang, M., Wu, Y., Zhao, T., Zhao, Z., Lu, C.: Pointsift: A sift-like network module for 3d point cloud semantic segmentation. arXiv preprint arXiv:1807.00652 (2018)
2018 arXiv
-
[35]
but should vqa expect them to? In: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition
Kervadec, C., Antipov, G., Baccouche, M., Wolf, C.: Roses are red, violets are blue... but should vqa expect them to? In: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition. pp. 2776–2785 (2021)
2021
-
[36]
arXiv preprint arXiv:1312.6114 (2013)
Kingma, D.P., Welling, M.: Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)
2013 arXiv
-
[37]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Kong, L., Liu, Y., Li, X., Chen, R., Zhang, W., Ren, J., Pan, L., Chen, K., Liu, Z.: Robo3d: Towards robust and reliable 3d perception against corruptions. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 19994–20006 (2023)
2023
-
[38]
In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops
Kortylewski, A., Egger, B., Schneider, A., Gerig, T., Morel-Forster, A., Vetter, T.: Empirically analyzing the effect of dataset biases on deep face recognition systems. In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops. pp. 2093–2102 (2018)
2018
-
[39]
Master’s thesis, University of Tront (2009)
Krizhevsky, A.: Learning multiple layers of features from tiny images. Master’s thesis, University of Tront (2009)
2009
-
[40]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Lai, X., Liu, J., Jiang, L., Wang, L., Zhao, H., Liu, S., Qi, X., Jia, J.: Stratified transformer for 3d point cloud segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8500–8509 (2022)
2022
-
[41]
Current Opinion in Behavioral Sciences29, 97–104 (2019)
Lake, B.M., Salakhutdinov, R., Tenenbaum, J.B.: The omniglot challenge: a 3-year progress report. Current Opinion in Behavioral Sciences29, 97–104 (2019)
2019
-
[42]
CS 231N7(7), 3 (2015)
Le, Y., Yang, X.: Tiny imagenet visual recognition challenge. CS 231N7(7), 3 (2015)
2015
-
[43]
Proceedings of the IEEE86(11), 2278–2324 (1998) BelHouse3D 17
LeCun, Y., Bottou, L., Bengio, Y., Haffner, P.: Gradient-based learning applied to document recognition. Proceedings of the IEEE86(11), 2278–2324 (1998) BelHouse3D 17
1998
-
[44]
In: British Machine Vision Conference (BMVC) (2018)
Li, W., Saeedi, S., McCormac, J., Clark, R., Tzoumanikas, D., Ye, Q., Huang, Y., Tang, R., Leutenegger, S.: Interiornet: Mega-scale multi-sensor photo-realistic indoor scenes dataset. In: British Machine Vision Conference (BMVC) (2018)
2018
-
[45]
arXiv e-prints pp
Li, Y., Li, S., Liu, X., Gong, M., Li, K., Chen, N., Wang, Z., Li, Z., Jiang, T., Yu, F., et al.: Sscbench: Monocular 3d semantic scene completion benchmark in street views. arXiv e-prints pp. arXiv–2306 (2023)
2023
-
[46]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Liang, W., Xue, F., Liu, Y., Zhong, G., Ming, A.: Unknown sniffer for object detec- tion: Don’t turn a blind eye to unknown objects. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3230–3239 (2023)
2023
-
[47]
IEEE Transactions on Pattern Analysis and Machine Intelligence 45(3), 3292–3310 (2022)
Liao, Y., Xie, J., Geiger, A.: Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d. IEEE Transactions on Pattern Analysis and Machine Intelligence 45(3), 3292–3310 (2022)
2022
-
[48]
In: Proceedings of the IEEE international conference on computer vision
Liu, Z., Luo, P., Wang, X., Tang, X.: Deep learning face attributes in the wild. In: Proceedings of the IEEE international conference on computer vision. pp. 3730– 3738 (2015)
2015
-
[49]
In: Proceedings eighth IEEE international conference on com- puter vision
Martin, D., Fowlkes, C., Tal, D., Malik, J.: A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In: Proceedings eighth IEEE international conference on com- puter vision. ICCV 2001. vol. 2...
2001
-
[50]
arXiv preprint arXiv:2211.09445 (2022)
Matuszka, T., Barton, I., Butykai, Á., Hajas, P., Kiss, D., Kovács, D., Kunsági- Máté, S., Lengyel, P., Németh, G., Pető, L., et al.: aimotive dataset: A multimodal dataset for robust autonomous driving with long-range perception. arXiv preprint arXiv:2211.09445 (2022)
2022 arXiv
-
[51]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Mo, K., Zhu, S., Chang, A.X., Yi, L., Tripathi, S., Guibas, L.J., Su, H.: Partnet: A large-scale benchmark for fine-grained and hierarchical part-level 3d object un- derstanding. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 909–918 (2019)
2019
-
[52]
arXiv preprint arXiv:2406.11835 (2024)
Nekrasov, A., Zhou, R., Ackermann, M., Hermans, A., Leibe, B., Rottmann, M.: Oodis: Anomaly instance segmentation benchmark. arXiv preprint arXiv:2406.11835 (2024)
2024 arXiv
-
[53]
In: 2008 Sixth Indian conference on computer vision, graphics & image processing
Nilsback, M.E., Zisserman, A.: Automated flower classification over a large number of classes. In: 2008 Sixth Indian conference on computer vision, graphics & image processing. pp. 722–729. IEEE (2008)
2008
-
[54]
In: International Conference on Med- ical Image Computing and Computer-Assisted Intervention
Özsoy, E., Örnek, E.P., Eck, U., Czempiel, T., Tombari, F., Navab, N.: 4d-or: Se- mantic scene graphs for or domain modeling. In: International Conference on Med- ical Image Computing and Computer-Assisted Intervention. pp. 475–485. Springer (2022)
2022
-
[55]
Pan, Y., Gao, B., Mei, J., Geng, S., Li, C., Zhao, H.: Semanticposs: A point cloud datasetwithlargequantityofdynamicinstances.In:2020IEEEIntelligentVehicles Symposium (IV). pp. 687–693. IEEE (2020)
2020
-
[56]
In: Proceedings of the IEEE/CVF international conference on computer vision
Peng, X., Bai, Q., Xia, X., Huang, Z., Saenko, K., Wang, B.: Moment matching for multi-source domain adaptation. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 1406–1415 (2019)
2019
-
[57]
Pointcept: Pointcept: A codebase for point cloud perception research.https:// github.com/Pointcept/Pointcept (2023)
2023
-
[58]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Qi, C.R., Su, H., Mo, K., Guibas, L.J.: Pointnet: Deep learning on point sets for 3d classification and segmentation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 652–660 (2017)
2017
-
[59]
Advances in neural information processing systems 30 (2017) 18 U
Qi, C.R., Yi, L., Su, H., Guibas, L.J.: Pointnet++: Deep hierarchical feature learn- ing on point sets in a metric space. Advances in neural information processing systems 30 (2017) 18 U. Raman Kumar et al
2017
-
[60]
In: Proceedings of the IEEE/CVF international conference on computer vision
Roberts, M., Ramapuram, J., Ranjan, A., Kumar, A., Bautista, M.A., Paczan, N., Webb, R., Susskind, J.M.: Hypersim: A photorealistic synthetic dataset for holis- tic indoor scene understanding. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 109...
2021
-
[61]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Ros, G., Sellart, L., Materzynska, J., Vazquez, D., Lopez, A.M.: The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3234–3243 (2016)
2016
-
[62]
The International Journal of Robotics Research37(6), 545–557 (2018)
Roynard, X., Deschaud, J.E., Goulette, F.: Paris-lille-3d: A large and high-quality ground-truth urban point cloud dataset for automatic segmentation and classifica- tion. The International Journal of Robotics Research37(6), 545–557 (2018)
2018
-
[63]
In: European Conference on Computer Vision
Rozenberszki, D., Litany, O., Dai, A.: Language-grounded indoor 3d semantic seg- mentation in the wild. In: European Conference on Computer Vision. pp. 125–141. Springer (2022)
2022
-
[64]
In: 2011 IEEE international conference on computer vision workshops (ICCV workshops)
Silberman, N., Fergus, R.: Indoor scene segmentation using a structured light sen- sor. In: 2011 IEEE international conference on computer vision workshops (ICCV workshops). pp. 601–608. IEEE (2011)
2011
-
[65]
Ad- vances in neural information processing systems30 (2017)
Snell, J., Swersky, K., Zemel, R.: Prototypical networks for few-shot learning. Ad- vances in neural information processing systems30 (2017)
2017
-
[66]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Song,S.,Lichtenberg,S.P.,Xiao,J.:Sunrgb-d:Argb-dsceneunderstandingbench- mark suite. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 567–576 (2015)
2015
-
[67]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., et al.: Scalability in perception for autonomous driving: Waymo open dataset. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognit...
2020
-
[68]
In: Proceedings of the IEEE/CVF international conference on computer vision
Thomas, H., Qi, C.R., Deschaud, J.E., Marcotegui, B., Goulette, F., Guibas, L.J.: Kpconv: Flexible and deformable convolution for point clouds. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 6411–6420 (2019)
2019
-
[69]
Advances in Neural Information Processing Systems34, 22221–22233 (2021)
Van Breugel, B., Kyono, T., Berrevoets, J., Van der Schaar, M.: Decaf: Generating fair synthetic data using causally-aware generative networks. Advances in Neural Information Processing Systems34, 22221–22233 (2021)
2021
-
[70]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops
Varney, N., Asari, V.K., Graehling, Q.: Dales: A large-scale aerial lidar data set for semantic segmentation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops. pp. 186–187 (2020)
2020
-
[71]
Wah, C., Branson, S., Welinder, P., Perona, P., Belongie, S.: The caltech-ucsd birds-200-2011 (cub-200-2011). Tech. Rep. CNS-TR-2011-001, California Institute of Technology (2011)
2011
-
[72]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Wang, H., Li, Z., Feng, L., Zhang, W.: Vim: Out-of-distribution with virtual-logit matching. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 4921–4930 (2022)
2022
-
[73]
ACM Transactions on Graphics (tog)38(5), 1–12 (2019)
Wang, Y., Sun, Y., Liu, Z., Sarma, S.E., Bronstein, M.M., Solomon, J.M.: Dynamic graph cnn for learning on point clouds. ACM Transactions on Graphics (tog)38(5), 1–12 (2019)
2019
-
[74]
Advances in Neural Information Processing Systems 35, 33330–33342 (2022)
Wu, X., Lao, Y., Jiang, L., Liu, X., Zhao, H.: Point transformer v2: Grouped vector attention and partition-based pooling. Advances in Neural Information Processing Systems 35, 33330–33342 (2022)
2022
-
[75]
In: 6th International Conference on Learning Repre- sentations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Workshop BelHouse3D 19 Track Proceedings
Wu, Y., Wu, Y., Gkioxari, G., Tian, Y.: Building generalizable agents with a real- istic and rich 3d environment. In: 6th International Conference on Learning Repre- sentations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Workshop BelHouse3D 19 Track Proceedings....
2018
-
[76]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., Xiao, J.: 3d shapenets: A deep representation for volumetric shapes. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1912–1920 (2015)
2015
-
[77]
In: Proceedings of the AAAI conference on artificial intelligence
Xiao, A., Huang, J., Guan, D., Zhan, F., Lu, S.: Transfer learning from synthetic to real lidar point cloud for semantic segmentation. In: Proceedings of the AAAI conference on artificial intelligence. vol. 36, pp. 2795–2803 (2022)
2022
-
[78]
arXiv preprint arXiv:1708.07747 (2017)
Xiao, H., Rasul, K., Vollgraf, R.: Fashion-mnist: a novel image dataset for bench- marking machine learning algorithms. arXiv preprint arXiv:1708.07747 (2017)
2017 arXiv
-
[79]
In: Proceedings of the IEEE international conference on computer vision
Xiao, J., Owens, A., Torralba, A.: Sun3d: A database of big spaces reconstructed using sfm and object labels. In: Proceedings of the IEEE international conference on computer vision. pp. 1625–1632 (2013)
2013
-
[80]
arXiv preprint arXiv:1802.06739 (2018)
Xie, L., Lin, K., Wang, S., Wang, F., Zhou, J.: Differentially private generative adversarial network. arXiv preprint arXiv:1802.06739 (2018)
2018 arXiv
-
[81]
In: Proceedings of the European conference on computer vision (ECCV)
Ye, X., Li, J., Huang, H., Du, L., Zhang, X.: 3d recurrent neural networks with con- text fusion for point cloud semantic segmentation. In: Proceedings of the European conference on computer vision (ECCV). pp. 403–417 (2018)
2018
-
[82]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Yeshwanth, C., Liu, Y.C., Nießner, M., Dai, A.: Scannet++: A high-fidelity dataset of 3d indoor scenes. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 12–22 (2023)
2023
-
[83]
IEEE journal of biomedical and health informatics24(8), 2378–2388 (2020)
Yoon, J., Drumright, L.N., Van Der Schaar, M.: Anonymization through data syn- thesis using generative adversarial networks (ads-gan). IEEE journal of biomedical and health informatics24(8), 2378–2388 (2020)
2020
-
[84]
Transactions of the Association for Computational Linguistics2, 67–78 (2014)
Young, P., Lai, A., Hodosh, M., Hockenmaier, J.: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics2, 67–78 (2014)
2014
-
[85]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition
Yu, F., Chen, H., Wang, X., Xian, W., Chen, Y., Liu, F., Madhavan, V., Darrell, T.: Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition. pp. 2636–2645 (2020)
2020
-
[86]
In: Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition
Zhang, X., He, Y., Xu, R., Yu, H., Shen, Z., Cui, P.: Nico++: Towards better benchmarking for domain generalization. In: Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition. pp. 16036–16047 (2023)
2023
-
[87]
arXiv preprint arXiv:2304.10266 (2023)
Zhao, B., Wang, J., Ma, W., Jesslen, A., Yang, S., Yu, S., Zendel, O., Theobalt, C., Yuille, A., Kortylewski, A.: Ood-cv-v2: An extended benchmark for robustness to out-of-distribution shifts of individual nuisances in natural images. arXiv preprint arXiv:2304.10266 (2023)
2023 arXiv
-
[88]
In: European conference on computer vision
Zhao, B., Yu, S., Ma, W., Yu, M., Mei, S., Wang, A., He, J., Yuille, A., Kortylewski, A.: Ood-cv: A benchmark for robustness to out-of-distribution shifts of individual nuisances in natural images. In: European conference on computer vision. pp. 163–
-
[89]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Zhao, H., Jiang, L., Fu, C.W., Jia, J.: Pointweb: Enhancing local neighborhood features for point cloud processing. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 5565–5573 (2019)
2019
-
[90]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Zhao, N., Chua, T.S., Lee, G.H.: Few-shot 3d point cloud semantic segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8873–8882 (2021)
2021
-
[91]
arXiv preprint arXiv:1907.12022 (2019) 20 U
Zhao, Z., Liu, M., Ramani, K.: Dar-net: Dynamic aggregation network for semantic scene segmentation. arXiv preprint arXiv:1907.12022 (2019) 20 U. Raman Kumar et al
2019 arXiv
-
[92]
In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IX 16
Zheng, J., Zhang, J., Li, J., Tang, R., Gao, S., Zhou, Z.: Structured3d: A large photo-realistic dataset for structured 3d modeling. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IX 16. pp. 519–535. Springer (2020)
2020
-
[93]
IEEE transactions on pattern analysis and machine intelligence 40(6), 1452–1464 (2017)
Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., Torralba, A.: Places: A 10 million image database for scene recognition. IEEE transactions on pattern analysis and machine intelligence 40(6), 1452–1464 (2017)
2017
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.