Pith. sign in

REVIEW 4 major objections 6 minor 162 references

Online Knowledge Integration for 3D Semantic Mapping: A Survey

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This survey claims that deep learning now embeds knowledge graphs and natural-language concepts directly into 3D semantic mapping pipelines, rather than adding them after mapping.

desk verdict Useful entry-point survey of knowledge integration in 3D semantic mapping, but the abstract's 'online' promise overreaches the surveyed methods; worth conditional acceptance after modest framing fixes. read the letter →

arxiv 2411.18147 v1 pith:LTSXU754 submitted 2024-11-27 cs.RO cs.CVcs.LG

classification cs.ROcs.CVcs.LG
keywords 3Dsemanticmappinglargelanguagemodelsvision-languagefoundationscenegraphsknowledgesegmentationrobotics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that prior knowledge—knowledge graphs, semantic scene graphs, and the common-sense knowledge in language models—can now be integrated directly into the sensor-data processing and semantic mapping pipeline, not merely bolted onto a finished geometric map. It reviews two complementary families of methods: those that build or predict 3D semantic scene graphs, and those that embed vision-language features or language-model reasoning into maps. The payoff for robotics, if the survey's reading is correct, is a map that can be queried in natural language, incrementally updated, and used for navigation, object search, task planning, and change detection. The paper's own tables show that 'online' is a spectrum, from real-time incremental graph construction to offline-distilled feature maps.

What carries the argument

The load-bearing machinery is the 3D semantic scene graph—nodes representing objects, rooms, and buildings, edges representing spatial or semantic relations—combined with a shared image-text embedding space supplied by vision-language foundation models and with language models as a source of relations and plans. The scene graph gives the map a hierarchical or flat symbolic structure; the embedding space makes that structure queryable by arbitrary text; and language models fill in relations, room types, or subgraphs for planning. The survey sorts methods by how the graph is generated (deterministic construction vs. learned prediction) and by how language features are attached (NeRF-based, geometry-based, or scene-graph-based), which is the classification that carries the survey's argument.

What would settle it

Run the surveyed methods on a streaming RGB-D benchmark where objects move and new rooms appear mid-run, and check on the robot's onboard compute whether the semantic layer (object labels, relations, language queries) updates within a bounded latency without retraining or cloud calls; if most methods fail this test, the paper's central 'online integration' claim is not borne out by its own evidence.

Watch

Extended reading notes

Core claim

The paper's central claim is that recent deep learning has collapsed the classical separation between geometric mapping, semantic labeling, and prior-knowledge reasoning into a single, knowledge-aware mapping process. Semantic scene graphs provide a symbolic backbone: nodes stand for physical entities and edges for spatial or semantic relations, and they can be constructed incrementally from streaming RGB-D data or predicted end-to-end by graph neural networks. Vision-language foundation models, such as the CLIP-style joint image-text embedding, let map elements be queried by arbitrary natural-language concepts rather than a fixed class set. Large language models and multimodal models contribute common-sense relations, room-type ontologies, and task plans that steer exploration. Together these mechanisms are claimed to make possible previously impossible applications, including language-specified object search and scalable task planning over large scene graphs.

Load-bearing premise

The survey's narrative assumes that the methods it groups under 'online integration' really are online—continuously updating the map from new sensor data—when several of its own cited methods require offline training, batch processing, or an internet connection.

Editorial extensions

If this is right

  • Robots will be able to query their maps with arbitrary natural-language object descriptions rather than a fixed list of classes.
  • Scene graphs can be maintained incrementally, so moving objects and room changes can be reflected in the map as new sensor data arrives.
  • Knowledge graphs and language models can inject common-sense relations at perception time, reducing the need for manually curated ontologies and fixed class sets.
  • Language-conditioned maps support downstream tasks such as object search, task planning, navigation, and change detection that geometric maps alone cannot express.
  • Combining scene graphs with vision-language features moves the map from point-wise labels to object-level, relation-aware representations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: a stricter definition of 'online'—continuous updates on the robot's own compute, with no retraining and no cloud calls—would remove several surveyed methods, so the field's true online capability is likely narrower than the abstract suggests.
  • Editorial extension: the same knowledge-injection mechanism could be applied to long-term map maintenance, letting a robot revise room labels and object relations as scenes change; the surveyed works mostly treat mapping as a one-shot build.
  • Editorial extension: a downstream-task benchmark—measuring navigation or object-search success rather than segmentation or predicate accuracy—would test whether language-integrated maps actually help robots, a comparison the survey does not provide.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper surveys recent methods for integrating prior knowledge into 3D semantic mapping, focusing on two families: semantic scene graphs (Section 3) and vision-language/large language models (Section 4). It reviews classical semantic mapping and knowledge representation (Section 2), then describes scene graph construction/prediction methods and their applications, and discusses VLFM-based open-vocabulary mapping and LLM-based reasoning, including comparative tables of methods and applications. The abstract frames the survey as a review of 'online integration of knowledge into semantic mapping' in which deep learning now allows 'full integration' of prior knowledge into sensor processing and mapping pipelines.

Significance. If the presentation is adjusted to accurately reflect the online/offline split, the survey provides a useful structured overview of recent work on scene-graph and language-model integration in 3D semantic mapping. Its strengths are the clear categorization into scene graphs and language models, the two overview tables, and the candid discussion of limitations in Sections 4.1.4 and 4.2.2. The survey does not claim new quantitative results, and the characterizations appear consistent with the cited literature at a high level. The central narrative, however, overstates the field's online maturity, and this mismatch must be resolved before the survey's scope claim is credible.

major comments (4)
  1. [Abstract; Section 4.1.1; Table 3] The abstract's claim of a 'focus on online integration' is not supported by the surveyed evidence, because the paper itself states that LERF and CLIP-Fields 'require extensive training for each scene' and 'support neither updates nor fine-tuning' (Section 4.1.1). Since these are presented as core VLFM integration approaches, the online focus needs to be either redefined to explicitly include offline scene-specific training, or narrowed to geometry-based/online methods, with the abstract adjusted accordingly.
  2. [Section 4.1.2; Table 3] OpenScene is described as distilling features into a 3D CNN (Section 4.1.2), which is an offline training step; nevertheless, it is listed in Table 3 without any online/offline indicator. The tables should systematically annotate each method's online/offline status (incremental updates, real-time suitability, onboard compute) so the reader can verify the survey's scope claim against the evidence.
  3. [Section 4.2.1; Table 4] SayPlan is explicitly said to 'requir[e] internet access and therefore not suitable for onboard-only deployment' (Section 4.2.1), yet it is presented under the survey's online-integration focus. This is a second clear case where the paper's own evidence contradicts the scope; the authors should either exclude such methods from the claimed focus or explicitly position them as off-board/cloud-based exceptions.
  4. [Abstract; Sections 4.1.4 and 4.2.2] The abstract's 'full integration' is too strong given the paper's own limitations list: VLFM-based maps 'lack any implicit specialized knowledge' and 'are incapable of reasoning except on a very basic level' (Section 4.1.4), and LLMs 'do not perform well at reasoning about new problems that are not in their training sets' (Section 4.2.2). Qualify 'full integration' to something like 'tight, in-pipeline integration for representative tasks' and reflect these limitations in the abstract.
minor comments (6)
  1. [Abstract] The phrase 'language models for respective capture of implicit common-sense knowledge' is ungrammatical; suggest 'language models for capturing implicit common-sense knowledge' or 'for the respective capture of implicit common-sense knowledge'.
  2. [Section 3.2] The D-SCG method is described as predicting unseen objects in 'incomplete ScanNet [88] scenes,' but reference [88] is the Matterport3D paper; the correct ScanNet reference is [101]. This citation error should be fixed.
  3. [Section 2.3] RDF and OWL are referred to as 'systems,' but they are standards or languages; the wording should be adjusted accordingly.
  4. [Section 3.4] The discussion of recall@k and mean recall@k would benefit from a brief definition of k and a citation to the original recall@k formulation, as the reader must currently infer the meaning.
  5. [Section 4.2.2] The citation 'cf. [160, 161, 162]' compacts multiple ethics/safety topics into a single spot; adding a short parenthetical or splitting the citations would improve readability.
  6. [Section 4.1 (general)] Given the claim of comprehensiveness, the absence of any discussion of open-vocabulary 3D Gaussian Splatting methods (e.g., LangSplat and similar works from 2023-2024) is noticeable; at least a brief mention would strengthen the VLFM section.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey makes no predictions or derivations and its conclusions rest on a broad set of independent external sources.

full rationale

This paper is a literature survey, so the circularity failure modes for derived claims, fitted parameters, and prediction-on-fit do not apply. The central claim—that deep learning now permits fuller integration of knowledge graphs and language concepts into semantic mapping pipelines—is a descriptive synthesis of the cited literature rather than a quantity derived from the paper's own inputs. The authors do cite their own prior work (e.g., Nüchter and Hertzberg [3] for the definition of a semantic map; SEMAP [29,30]; Günther et al. [10,42]; Wiemann et al. [8]), but these citations are background or contextual and none carries the survey's conclusions. The survey's own admission that some included methods (LERF, CLIP-Fields, OpenScene, SayPlan) are offline or require internet access is a potential scope or correctness concern about the 'online integration' emphasis, not a circularity: it does not convert any input into an output by construction. Because the paper is externally grounded in a wide range of independent references and makes no predictions, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey's synthesis rests on the trustworthiness of the cited papers, on the Nüchter-Hertzberg scope definition, and on the assumption that filtered LLM output can serve as knowledge for robotics. These are domain assumptions rather than fitted parameters; no free parameters or invented entities appear.

assumptions (3)
  • domain assumption The surveyed primary papers accurately report capabilities, datasets, and quantitative improvements.
    The survey does not re-run any method; its synthesis inherits the source papers' reported numbers, such as SGFormer improvements on 3DSSG or SayPlan's 73% accuracy.
  • domain assumption The Nüchter-Hertzberg definition of semantic mapping delimits what counts as in-scope for the review.
    Section 1 adopts this definition as the survey's organizing frame, so methods outside it, such as purely 2D scene graphs in image captioning, are excluded.
  • domain assumption Language-model output, after filtering, is reliable enough to produce ontologies and scene graph relations.
    Section 4 relies on LLM/LMM outputs for room classification (Strader et al.), relation prediction (ConceptGraphs), and plan generation (SayPlan), while Section 4.2.2 concedes hallucination and out-of-distribution reasoning failures.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Online Knowledge Integration for 3D Semantic Mapping: A Survey." pith.science (2026). https://pith.science/paper/LTSXU754

@misc{pith2026241118147,
  author       = {Pith},
  title        = {Pith review of: Online Knowledge Integration for 3D Semantic Mapping: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LTSXU754}},
  note         = {Machine review of arXiv:2411.18147}
}
read the original abstract

Semantic mapping is a key component of robots operating in and interacting with objects in structured environments. Traditionally, geometric and knowledge representations within a semantic map have only been loosely integrated. However, recent advances in deep learning now allow full integration of prior knowledge, represented as knowledge graphs or language concepts, into sensor data processing and semantic mapping pipelines. Semantic scene graphs and language models enable modern semantic mapping approaches to incorporate graph-based prior knowledge or to leverage the rich information in human language both during and after the mapping process. This has sparked substantial advances in semantic mapping, leading to previously impossible novel applications. This survey reviews these recent developments comprehensively, with a focus on online integration of knowledge into semantic mapping. We specifically focus on methods using semantic scene graphs for integrating symbolic prior knowledge and language models for respective capture of implicit common-sense knowledge and natural language concepts

Figures

Figures reproduced from arXiv: 2411.18147 by the authors.

Figure 1
Figure 1. Hierarchical scene graph from [35] were still mainly geometric, namely left/right, in front of/behind, bigger/smaller, occlusion relationships, and relationships to other layers, but this framework pro￾vided a key foundation for later advances. Improving on the idea of hierarchical scene graphs, Rosinol et al. [67] presented a method for automatic 3D scene graph construction from visual-inertial data. The described … view at source ↗
Figure 2
Figure 2. Flat incremental scene graph from [79] Qiu and Christensen [28] integrate a knowledge graph based on WordNet, ConceptNet, and Visual Genome which models pairwise relationships. Us￾ing a Knowledge-Scene Graph Network (KSGN) based on the Graph Bridging Network [85], the knowledge graph is included in the message passing process in the scene graph, improving scene graph prediction on the 3DSSG160 dataset. In addition t… view at source ↗
Figure 3
Figure 3. Association of textual input with visual concepts and properties in a single embedding space using the CLIP model. [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Relevancy maps as provided by LERF [130] in the ”kitchen” [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Result segments for the query ”chair” from a map generated [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: Example for Embodied Question Answering using an LLM [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

162 extracted references · 20 canonical work pages

  1. [1]

    Stachniss, J

    C. Stachniss, J. J. Leonard, S. Thrun, Simultaneous local- ization and mapping, in: B. Siciliano, O. Khatib (Eds.), Springer Handbook of Robotics, Springer, 2016, pp. 1153–

  2. [2]

    Cadena, L

    C. Cadena, L. Carlone, H. Carrillo, Y . Latif, D. Scaramuzza, J. Neira, I. Reid, J. J. Leonard, Past, present, and future of Simultaneous Localization and Mapping: Toward the robust- perception age, IEEE Trans. Robotics 32 (2016) 1309–1332. doi:10.1109/TRO.2016.2624754

  3. [3]

    N ¨uchter, J

    A. N ¨uchter, J. Hertzberg, Towards semantic maps for mo- bile robots, Robotics Auton. Syst. 56 (2008) 915–926. doi:10.1016/j.robot.2008.08.001

  4. [4]

    X. Han, S. Li, X. Wang, W. Zhou, Semantic mapping for mobile robots in indoor scenes: A survey, Inf. 12 (2021) 92. doi:10.3390/INFO12020092

  5. [5]

    Gouidis, A

    F. Gouidis, A. Vassiliades, T. Patkos, A. Argyros, N. Bassili- ades, D. Plexousakis, A review on intelligent object perception methods combining knowledge-based reasoning and machine learning, CEUR Workshop Proc. 2600 (2019)

  6. [6]

    Grisetti, C

    G. Grisetti, C. Stachniss, W. Burgard, Improved techniques for grid mapping with rao-blackwellized particle filters, IEEE Trans. Robotics 23 (2007) 34–46

  7. [7]

    N ¨uchter, K

    A. N ¨uchter, K. Lingemann, J. Hertzberg, H. Surmann, 6D SLAM – 3D mapping outdoor environments, J. Field Robot. 24 (2007) 699–722. doi:10.1002/rob.20209

  8. [8]

    Wiemann, I

    T. Wiemann, I. Mitschke, A. Mock, J. Hertzberg, Surface reconstruction from arbitrarily large point clouds, in: 2018 Second IEEE International Conference on Robotic Computing (IRC), IEEE, 2018, pp. 278–281

Show all 162 references
  1. [9]

    P ¨utz, T

    S. P ¨utz, T. Wiemann, J. Sprickerhof, J. Hertzberg, 3D nav- igation mesh generation for path planning in uneven terrain, IFAC-PapersOnLine 49 (2016) 212–217

  2. [10]

    G ¨unther, T

    M. G ¨unther, T. Wiemann, S. Albrecht, J. Hertzberg, Model- based furniture recognition for building semantic object maps, Artif. Intell. 247 (2017) 336–351. doi: 10.1016/ J.ARTINT.2014.12.007

  3. [11]

    Hornung, K

    A. Hornung, K. M. Wurm, M. Bennewitz, C. Stachniss, W. Burgard, OctoMap: an e fficient probabilistic 3D mapping framework based on octrees, Auton. Robots 34 (2013) 189–

  4. [12]

    Zhang, S

    J. Zhang, S. Singh, LOAM: lidar odometry and mapping in real-time, in: Robotics: Science and Systems X (RSS 2014),

  5. [13]

    T. Shan, B. J. Englot, LeGO-LOAM: Lightweight and ground- optimized lidar odometry and mapping on variable terrain, in: 2018 IEEE /RSJ International Conference on Intelligent Robots and Systems (IROS 2018), IEEE, 2018, pp. 4758–

  6. [14]

    T. Shan, B. J. Englot, D. Meyers, W. Wang, C. Ratti, D. Rus, LIO-SAM: tightly-coupled lidar inertial odometry via smooth- ing and mapping, in: IEEE /RSJ International Conference on Intelligent Robots and Systems (IROS 2020), IEEE, 2020, pp. 5135–5142. doi:10.1109/IROS45743.202...

  7. [15]

    Vizzo, T

    I. Vizzo, T. Guadagnino, B. Mersch, L. Wiesmann, J. Behley, C. Stachniss, KISS-ICP: in defense of point-to-point ICP – simple, accurate, and robust registration if done the right way, IEEE Robotics Autom. Lett. 8 (2023) 1029–1036. doi:10.1109/LRA.2023.3236571

  8. [16]

    Izadi, R

    S. Izadi, R. A. Newcombe, D. Kim, O. Hilliges, D. Molyneaux, S. Hodges, P. Kohli, J. Shotton, A. J. Davison, A. W. Fitzgib- bon, KinectFusion: real-time dynamic 3D surface reconstruc- tion and interaction, in: International Conference on Com- puter Graphics and Interactive Tec...

  9. [17]

    Whelan, M

    T. Whelan, M. Kaess, M. Fallon, H. Johannsson, J. Leonard, J. McDonald, Kintinuous: Spatially Extended KinectFusion, Technical Report MIT-CSAIL-TR-2012-020, CSAIL Techni- cal Reports, 2012

  10. [18]

    Whelan, S

    T. Whelan, S. Leutenegger, R. F. Salas-Moreno, B. Glocker, A. J. Davison, ElasticFusion: Dense SLAM without a pose graph, in: Robotics: science and systems, volume 11, Rome, Italy, 2015, p. 3

  11. [19]

    Campos, R

    C. Campos, R. Elvira, J. J. G. Rodr ´ıguez, J. M. M. Mon- tiel, J. D. Tard ´os, ORB-SLAM3: An accurate open- source library for visual, visual-inertial, and multimap SLAM, IEEE Trans. Robotics 37 (2021) 1874–1890. doi: 10.1109/ TRO.2021.3075644

  12. [20]

    Engel, V

    J. Engel, V . Koltun, D. Cremers, Direct sparse odometry, IEEE Trans. Pattern Anal. Mach. Intell. 40 (2018) 611–625. doi:10.1109/TPAMI.2017.2658577

  13. [21]

    T. Qin, P. Li, S. Shen, Vins-mono: A robust and versatile monocular visual-inertial state estima- tor, IEEE Trans. Robotics 34 (2018) 1004–1020. doi:10.1109/TRO.2018.2853729

  14. [22]

    T. Qin, S. Shen, Online temporal calibration for monocu- lar visual-inertial systems, in: 2018 IEEE /RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2018, pp. 3662–3669

  15. [23]

    Mildenhall, P

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Bar- ron, R. Ramamoorthi, R. Ng, NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, CoRR abs/2003.08934 (2020). doi: 10.48550/arXiv.2003.08934. arXiv:2003.08934

  16. [24]

    Rosinol, J

    A. Rosinol, J. J. Leonard, L. Carlone, NeRF-SLAM: Real-time dense monocular slam with neural radiance fields, in: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2023, pp. 3437–3444

  17. [25]

    Kerbl, G

    B. Kerbl, G. Kopanas, T. Leimk¨uhler, G. Drettakis, 3d gaussian splatting for real-time radiance field rendering., ACM Trans. Graph. 42 (2023) 139–1

  18. [26]

    Matsuki, R

    H. Matsuki, R. Murai, P. H. Kelly, A. J. Davison, Gaussian splatting slam, in: Proceedings of the IEEE /CVF Confer- ence on Computer Vision and Pattern Recognition, 2024, pp. 18039–18048

  19. [27]

    C. Yan, D. Qu, D. Xu, B. Zhao, Z. Wang, D. Wang, X. Li, GS-SLAM: dense visual SLAM with 3D gaussian splatting, in: IEEE /CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), IEEE, 2024, pp. 19595–19604. doi:10.1109/CVPR52733.2024.01853

  20. [28]

    Y . Qiu, H. I. Christensen, 3D scene graph prediction on point clouds using knowledge graphs, in: 19 th IEEE In- ternational Conference on Automation Science and Engi- neering (CASE 2023), IEEE, 2023, pp. 1–7. doi: 10.1109/ CASE56687.2023.10260650

  21. [29]

    Deeken, T

    H. Deeken, T. Wiemann, K. Lingemann, J. Hertzberg, SEMAP – a semantic environment mapping framework, in: 2015 Eu- ropean Conference on Mobile Robots (ECMR 2015), IEEE, 2015, pp. 1–6. doi:10.1109/ECMR.2015.7324176

  22. [30]

    Deeken, T

    H. Deeken, T. Wiemann, J. Hertzberg, Grounding semantic maps in spatial databases, Robotics Auton. Syst. 105 (2018) 146–165. doi:10.1016/J.ROBOT.2018.03.011

  23. [32]

    Zhang, D

    C. Zhang, D. Han, Y . Qiao, J. U. Kim, S. Bae, S. Lee, C. S. Hong, Faster segment anything: Towards lightweight SAM for mobile applications, CoRR abs /2306.14289 (2023). doi:10.48550/ARXIV.2306.14289. arXiv:2306.14289

  24. [33]

    X. Zhao, W. Ding, Y . An, Y . Du, T. Yu, M. Li, M. Tang, J. Wang, Fast segment anything, CoRR abs/2306.12156 (2023). doi: 10.48550/ARXIV.2306.12156. arXiv:2306.12156

  25. [34]

    N. Ravi, V . Gabeur, Y . Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R ¨adle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V . Alwala, N. Carion, C. Wu, R. B. Girshick, P. Doll ´ar, C. Feichtenhofer, SAM 2: Segment anything in images and videos, CoRR abs /2408.00714 (2024). do...

  26. [35]

    Rosinol, A

    A. Rosinol, A. Violette, M. Abate, N. Hughes, Y . Chang, J. Shi, A. Gupta, L. Carlone, Kimera: from SLAM to spatial percep- tion with 3D dynamic scene graphs, Int. J. Rob. Res. (2021)

  27. [36]

    Grinvald, F

    M. Grinvald, F. Furrer, T. Novkovic, J. J. Chung, C. Cadena, R. Siegwart, J. I. Nieto, V olumetric instance-aware semantic mapping and 3D object discovery, IEEE Robotics Autom. Lett. 4 (2019) 3037–3044. doi:10.1109/LRA.2019.2923960

  28. [37]

    Freda, PLVS: A SLAM system with points, lines, volu- metric mapping, and 3D incremental segmentation, CoRR abs/2309.10896 (2023)

    L. Freda, PLVS: A SLAM system with points, lines, volu- metric mapping, and 3D incremental segmentation, CoRR abs/2309.10896 (2023). doi: 10.48550/ARXIV.2309.10896. arXiv:2309.10896

  29. [38]

    Banegas-Luna, J.-R

    A.-J. Banegas-Luna, J.-R. Ruiz-Sarmiento, G. Ambrosio- Cestero, J. Gonz ´alez-Jim´enez, Towards a voxelized semantic representation of the workspace of mobile robots, in: I. Rojas, G. Joya, A. Catal`a (Eds.), 17th International Work-Conference on Artificial Neural Networks (IW...

  30. [39]

    Fern ´andez-Chaves, J

    D. Fern ´andez-Chaves, J. R. Ruiz-Sarmiento, N. Petkov, J. Gonz ´alez-Jim´enez, ViMantic, a distributed robotic ar- chitecture for semantic mapping in indoor environments, Knowl. Based Syst. 232 (2021) 107440. doi: 10.1016/ J.KNOSYS.2021.107440

  31. [40]

    Schmid, M

    L. Schmid, M. Abate, Y . Chang, L. Carlone, Khronos: A uni- fied approach for spatio-temporal metric-semantic slam in dy- namic environments, in: Proc. of Robotics: Science and Sys- tems (RSS), 2024. doi:10.15607/RSS.2024.XX.081

  32. [41]

    Coradeschi, A

    S. Coradeschi, A. Sa ffiotti, An introduction to the anchoring problem, Robot. Auton. Syst. 43 (2003) 85–96

  33. [42]

    G ¨unther, J

    M. G ¨unther, J. R. Ruiz-Sarmiento, C. Galindo, J. Gonz ´alez- Jim´enez, J. Hertzberg, Context-aware 3D object anchoring for mobile robots, Robot. Auton. Syst. 110 (2018) 12–32. doi:10.1016/j.robot.2018.08.016

  34. [43]

    M. Feng, H. Hou, L. Zhang, Z. Wu, Y . Guo, A. Mian, 3D spa- tial multimodal knowledge accumulation for scene graph pre- diction in point cloud, Proc. IEEE Comput. Soc. Conf. Com- put. Vis. Pattern Recognit. (2023) 9182–9191

  35. [44]

    Zhang, S

    S. Zhang, S. Li, A. Hao, H. Qin, Knowledge-inspired 3D scene graph prediction in point cloud, Adv. Neural Inf. Process. Syst. (2021) 18620–18632

  36. [45]

    D. Lang, D. Paulus, Semantic maps for robotics, in: Proc. Workshop Workshop AI Robot.(ICRA), 2014, pp. 14–18

  37. [46]

    Tenorth, M

    M. Tenorth, M. Beetz, KnowRob: A knowledge processing infrastructure for cognition-enabled robots, I. J. Robotics Res. 32 (2013) 566–590. doi:10.1177/0278364913481635

  38. [47]

    M. A. Cornejo-Lupa, R. P. Ticona-Herrera, Y . Cardinale, D. Barrios-Aranibar, A Survey of Ontologies for Simultaneous Localization and Mapping in Mobile Robots, ACM Comput. Surv. 53 (2020) 1–26. doi:10.1145/3408316

  39. [48]

    H. Wang, J. Ren, A semantic map for indoor robot navigation based on predicate logic, Int. J. Knowl. Syst. Sci. 11 (2020) 1–21. doi:10.4018/IJKSS.2020010101

  40. [49]

    Z. Liu, G. von Wichert, A generalizable knowledge frame- work for semantic indoor mapping based on markov logic net- works and data driven MCMC, Future Gener. Comput. Syst. 36 (2014) 42–56. doi:10.1016/J.FUTURE.2013.06.026

  41. [50]

    M. A. Cornejo-Lupa, Y . Cardinale, R. Ticona-Herrera, D. Barrios-Aranibar, M. Andrade, J. Diaz-Amado, On- toSLAM: An Ontology for Representing Location and Si- multaneous Mapping Information for Autonomous Robots, Robotics 10 (2021) 125. doi:10.3390/robotics10040125

  42. [51]

    Kalanat, A

    N. Kalanat, A. Kovashka, Symbolic image detection using scene and knowledge graphs, CoRR abs /2206.04863 (2022). doi:10.48550/ARXIV.2206.04863. arXiv:2206.04863

  43. [52]

    Doing, R

    V . Doing, R. Wisnesky, Towards a more reasonable se- mantic web, CoRR abs /2407.19095 (2024). doi: 10.48550/ ARXIV.2407.19095. arXiv:2407.19095

  44. [53]

    S. Ji, S. Pan, E. Cambria, P. Marttinen, P. S. Yu, A survey on knowledge graphs: Representation, acquisition, and appli- cations, IEEE Trans. Neural Networks Learn. Syst. 33 (2022) 494–514. doi:10.1109/TNNLS.2021.3070843

  45. [54]

    Ferrucci, E

    D. Ferrucci, E. Brown, J. Chu-Carroll, J. Fan, D. Gondek, A. A. Kalyanpur, A. Lally, J. W. Murdock, E. Nyberg, J. Prager, N. Schlaefer, C. Welty, Building Watson: An Overview of the DeepQA Project, AI Mag. 31 (2010) 59–79. doi:10.1609/ aimag.v31i3.2303

  46. [55]

    G. A. Miller, WordNet: A lexical database for English, Com- mun. ACM 38 (1995) 39–41. doi:10.1145/219717.219748

  47. [56]

    Krishna, Y

    R. Krishna, Y . Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y . Kalantidis, L.-J. Li, D. A. Shamma, M. S. Bern- stein, L. Fei-Fei, Visual genome: Connecting language and vision using crowdsourced dense image annotations, Int. J. Comput. Vis. 123 (2017) 32–73

  48. [57]

    Speer, J

    R. Speer, J. Chin, C. Havasi, ConceptNet 5.5: An open multi- lingual graph of general knowledge, Proc. Conf. AAAI Artif. Intell. 31 (2017)

  49. [58]

    X. Wang, Y . Ye, A. Gupta, Zero-shot recognition via semantic embeddings and knowledge graphs, in: 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2018), Computer Vision Foundation / IEEE Computer Society, 2018, pp. 6857–6866. doi:10.1109/CVPR.2018.00717

  50. [59]

    C.-W. Lee, W. Fang, C.-K. Yeh, Y .-C. F. Wang, Multi-label Zero-Shot Learning with Structured Knowledge Graphs, in: 2018 IEEE /CVF Conference on Computer Vision and Pat- tern Recognition, IEEE, 2018, pp. 1576–1585. doi: 10.1109/ CVPR.2018.00170

  51. [60]

    Beetz, D

    M. Beetz, D. Beßler, A. Haidu, M. Pomarlan, A. K. Bozcuoglu, G. Bartels, KnowRob 2.0 – A 2 nd generation knowledge processing framework for cognition-enabled robotic agents, in: 2018 IEEE International Conference on Robotics and Automation (ICRA 2018), 2018, pp. 512–519. doi: ...

  52. [61]

    N. F. Noy, M. Klein, Ontology evolution: Not the same as schema evolution, Knowl. Inf. Syst. 6 (2004) 428–440

  53. [62]

    Kamp ffmeyer, Y

    M. Kamp ffmeyer, Y . Chen, X. Liang, H. Wang, Y . Zhang, E. P. Xing, Rethinking Knowledge Graph Propagation for Zero- Shot Learning, in: 2019 IEEE /CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), IEEE, 2019, pp. 11479–11488. doi:10.1109/CVPR.2019.01175

  54. [63]

    P. S. H. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. K ¨uttler, M. Lewis, W. Yih, T. Rockt ¨aschel, S. Riedel, D. Kiela, Retrieval-augmented generation for knowledge-intensive NLP tasks, in: Advances in Neural Infor- mation Processing Systems 33: Annual...

  55. [64]

    Strader, N

    J. Strader, N. Hughes, W. Chen, A. Speranzon, L. Carlone, Indoor and outdoor 3D scene graph generation via language- enabled spatial ontologies, IEEE Robotics Autom. Lett. 9 (2024) 4886–4893. doi:10.1109/LRA.2024.3384084

  56. [65]

    Bosselut, H

    A. Bosselut, H. Rashkin, M. Sap, C. Malaviya, A. Celikyilmaz, Y . Choi, COMET: commonsense transformers for automatic knowledge graph construction, in: Proceedings of the 57 th Conference of the Association for Computational Linguistics (ACL 2019), V olume 1: Long Papers, Asso...

  57. [66]

    Armeni, Z

    I. Armeni, Z. He, A. Zamir, J. Gwak, J. Malik, M. Fischer, S. Savarese, 3D scene graph: A structure for unified seman- tics, 3D space, and camera, in: 2019 IEEE /CVF International Conference on Computer Vision (ICCV 2019), IEEE, 2019, pp. 5663–5672. doi:10.1109/ICCV.2019.00576

  58. [67]

    Rosinol, A

    A. Rosinol, A. Gupta, M. Abate, J. Shi, L. Carlone, 3D dy- namic scene graphs: Actionable spatial perception with places, objects, and humans, Robot. Sci. Syst. (2020)

  59. [68]

    Hughes, Y

    N. Hughes, Y . Chang, L. Carlone, Hydra: A real-time spatial perception system for 3D scene graph construction and opti- mization, in: Robotics: Science and Systems XVIII, 2022. doi:10.15607/RSS.2022.XVIII.050

  60. [69]

    Kim, J.-M

    U.-H. Kim, J.-M. Park, T.-J. Song, J.-H. Kim, 3-D scene graph: A sparse and semantic representation of physical environments for intelligent agents, IEEE Trans. Cybern. 50 (2020) 4921– 4933

  61. [70]

    J. Wald, H. Dhamo, N. Navab, F. Tombari, Learning 3D se- mantic scene graphs from 3D indoor reconstructions, 2020

  62. [71]

    Kabalar, S.-C

    J. Kabalar, S.-C. Wu, J. Wald, K. Tateno, N. Navab, F. Tombari, Towards long-term retrieval-based visual localization in indoor environments with changes, IEEE Robotics Autom. Lett. 8 (2023) 1975–1982

  63. [72]

    J. Bae, D. Shin, K. Ko, J. Lee, U. Kim, A survey on 3D scene graphs: Definition, generation and application, in: 10 th International Conference on Robot Intelligence Technology and Applications (RiTA 2022), Springer, 2022, pp. 136–147. doi:10.1007/978-3-031-26889-2\ 13

  64. [73]

    Teney, L

    D. Teney, L. Liu, A. Van Den Hengel, Graph-structured repre- sentations for visual question answering, in: 2017 IEEE Con- ference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 3233–3241

  65. [74]

    T. Yao, Y . Pan, Y . Li, T. Mei, Exploring visual relationship for image captioning, in: European Conference on Computer Vision (ECCV 2018), Springer International Publishing, 2018, pp. 711–727

  66. [75]

    C. Agia, K. M. Jatavallabhula, M. Khodeir, O. Miksik, V . Vi- neet, M. Mukadam, L. Paull, F. Shkurti, Taskography: Evalu- ating robot task planning over large 3D scene graphs, in: Pro- ceedings of the 5th Conference on Robot Learning, volume 164 of Proceedings of Machine Learn...

  67. [76]

    J. Gu, H. Zhao, Z. Lin, S. Li, J. Cai, M. Ling, Scene graph gen- eration with external knowledge and image reconstruction, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 1969–1978

  68. [77]

    K. He, G. Gkioxari, P. Dollar, R. Girshick, Mask R-CNN, in: Proceedings of the IEEE International Conference on Com- puter Vision (ICCV), 2017

  69. [78]

    J. Wald, A. Avetisyan, N. Navab, F. Tombari, M. Nießner, RIO: 3D object instance re-localization in changing indoor environments, in: 2019 IEEE /CVF International Confer- ence on Computer Vision (ICCV 2019), 2019, pp. 7657–7666. doi:10.1109/ICCV.2019.00775

  70. [79]

    S.-C. Wu, J. Wald, K. Tateno, N. Navab, F. Tombari, Scene- GraphFusion: Incremental 3D scene graph prediction from RGB-D sequences, 2021. doi: 10.48550/arXiv.2103.14898. arXiv:2103.14898

  71. [80]

    X. Li, D. Guo, H. Liu, F. Sun, Embodied semantic scene graph generation, in: Proceedings of the 5 th Conference on Robot Learning, volume 164 of Proceedings of Machine Learning Research, PMLR, 2022, pp. 1585–1594

  72. [81]

    C. Qi, J. Yin, Z. Zhang, J. Tang, Dynamic scene graph gen- eration of point clouds with structural representation learning, Tsinghua Sci. Technol. 29 (2024) 232–243

  73. [82]

    K. Cho, B. van Merri ¨enboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, Y . Bengio, Learning phrase repre- sentations using RNN encoder–decoder for statistical machine translation, in: Proceedings of the 2014 Conference on Empir- ical Methods in Natural Language Proce...

  74. [83]

    C. R. Qi, H. Su, K. Mo, L. J. Guibas, PointNet: Deep Learning on point sets for 3D classification and segmentation, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), 2017, pp. 77–85. doi:10.1109/CVPR.2017.16

  75. [84]

    S.-C. Wu, K. Tateno, N. Navab, F. Tombari, Incremental 3D semantic scene graph prediction from RGB sequences, Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recog- nit. (2023) 5064–5074

  76. [85]

    Zareian, S

    A. Zareian, S. Karaman, S.-F. Chang, Bridging knowledge graphs to generate scene graphs, in: European Conference on Computer Vision (ECCV 2020), Springer-Verlag, 2020, pp. 606–623

  77. [86]

    C. Lv, M. Qi, X. Li, Z. Yang, H. Ma, SGFormer: Semantic graph transformer for point cloud-based 3D scene graph gen- eration, Proc. Conf. AAAI Artif. Intell. 38 (2024) 4035–4043

  78. [87]

    Badreddine, A

    S. Badreddine, A. d’Avila Garcez, L. Serafini, M. Spranger, Logic tensor networks, Artif. Intell. 303 (2020)

  79. [88]

    A. X. Chang, A. Dai, T. A. Funkhouser, M. Halber, M. Nießner, M. Savva, S. Song, A. Zeng, Y . Zhang, Matterport3D: Learning from RGB-D data in indoor environments, in: 2017 International Conference on 3D Vision (3DV 2017), IEEE Computer Society, 2017, pp. 667–676. doi: 10.1109...

  80. [89]

    Giuliari, G

    F. Giuliari, G. Skenderi, M. Cristani, A. D. Bue, Y . Wang, Leveraging commonsense for object localisation in partial scenes, IEEE Trans. Pattern Anal. Mach. Intell. 45 (2023) 12038–12049

  81. [90]

    C. Han, H. Li, J. Xu, B. Dong, Y . Wang, X. Zhou, S. Zhao, Unbiased 3D semantic scene graph prediction in point cloud using deep learning, Appl. Sci. 13 (2023) 5657. doi: 10.3390/ app13095657

  82. [91]

    Ravichandran, L

    Z. Ravichandran, L. Peng, N. Hughes, J. D. Gri ffith, L. Car- lone, Hierarchical representations and explicit memory: Learn- ing e ffective navigation policies on 3D scene graphs using graph neural networks, in: 2022 International Conference on Robotics and Automation (ICRA 20...

  83. [92]

    S ¨underhauf, Where are the keys? – Learning object- centric navigation policies on semantic maps with graph convolutional networks, CoRR abs /1909.07376 (2019)

    N. S ¨underhauf, Where are the keys? – Learning object- centric navigation policies on semantic maps with graph convolutional networks, CoRR abs /1909.07376 (2019). arXiv:1909.07376

  84. [93]

    Lingelbach, C

    M. Lingelbach, C. Li, M. Hwang, A. Kurenkov, A. Lou, R. Mart´ın-Mart´ın, R. Zhang, L. Fei-Fei, J. Wu, Task-driven graph attention for hierarchical relational object navigation, in: IEEE International Conference on Robotics and Automa- tion (ICRA 2023), IEEE, 2023, pp. 886–893....

  85. [94]

    Z. Hu, Y . Dong, K. Wang, Y . Sun, Heterogeneous graph transformer, in: WWW ’20: The Web Conference 2020, ACM / IW3C2, 2020, pp. 2704–2710. doi: 10.1145/ 3366423.3380027

  86. [95]

    Amiri, K

    S. Amiri, K. Chandan, S. Zhang, Reasoning with scene graphs for robot planning under partial observability, IEEE Robotics Autom. Lett. 7 (2022) 5560–5567. doi: 10.1109/ LRA.2022.3157567

  87. [96]

    Looper, J

    S. Looper, J. Rodriguez-Puigvert, R. Siegwart, C. Cadena, L. Schmid, 3D VSG: Long-term semantic scene change pre- diction through 3D variable scene graphs, in: 2023 IEEE In- ternational Conference on Robotics and Automation (ICRA), IEEE, 2023, pp. 8179–8186

  88. [97]

    Chang, L

    Y . Chang, L. Ballotta, L. Carlone, D-Lite: Navigation- oriented compression of 3D scene graphs for multi-robot col- laboration, IEEE Robotics Autom. Lett. 8 (2023) 7527–7534. doi:10.1109/LRA.2023.3320011

  89. [98]

    F. Xia, A. R. Zamir, Z. He, A. Sax, J. Malik, S. Savarese, Gibson Env: Real-world perception for embodied agents, in: 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2018), Computer Vision Foundation / IEEE Computer Society, 2018, pp. 9068–9079. doi: 10.1...

  90. [99]

    Rosinol, A

    A. Rosinol, A. Gupta, M. Abate, J. Shi, L. Car- lone, uHumans dataset, https://web.mit.edu/sparklab/ datasets/uHumans/, 2020. Accessed: 2024-9-22

  91. [100]

    C. Li, F. Xia, R. Mart ´ın-Mart´ın, M. Lingelbach, S. Srivas- tava, B. Shen, K. E. Vainio, C. Gokmen, G. Dharan, T. Jain, A. Kurenkov, C. K. Liu, H. Gweon, J. Wu, L. Fei-Fei, S. Savarese, iGibson 2.0: Object-centric simulation for robot learning of everyday household tasks, in...

  92. [101]

    A. Dai, A. X. Chang, M. Savva, M. Halber, T. A. Funkhouser, M. Nießner, ScanNet: Richly-annotated 3D reconstructions of indoor scenes, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), IEEE Computer Society, 2017, pp. 2432–2443. doi:10.1109/CVPR.2017.261

  93. [102]

    Armeni, O

    I. Armeni, O. Sener, A. R. Zamir, H. Jiang, I. K. Brilakis, M. Fischer, S. Savarese, 3D semantic parsing of large-scale indoor spaces, in: 2016 IEEE Conference on Computer Vi- sion and Pattern Recognition (CVPR 2016), IEEE Computer Society, 2016, pp. 1534–1543. doi:10.1109/CVP...

  94. [103]

    Roynard, J

    X. Roynard, J. Deschaud, F. Goulette, Paris-Lille-3D: A large and high-quality ground-truth urban point cloud dataset for au- tomatic segmentation and classification, Int. J. Robotics Res. 37 (2018) 545–557. doi:10.1177/0278364918767506

  95. [104]

    Kolve, R

    E. Kolve, R. Mottaghi, D. Gordon, Y . Zhu, A. Gupta, A. Farhadi, AI2-THOR: an interactive 3D environment for vi- sual AI, CoRR abs/1712.05474 (2017). arXiv:1712.05474

  96. [105]

    J. Wald, T. Sattler, S. Golodetz, T. Cavallari, F. Tombari, Be- yond controlled environments: 3D camera re-localization in changing indoor scenes, in: European Conference on Com- puter Vision (ECCV 2020), Springer, 2020, pp. 467–487. doi:10.1007/978-3-030-58571-6\ 28

  97. [106]

    C. Lu, R. Krishna, M. S. Bernstein, L. Fei-Fei, Visual relation- ship detection with language priors, in: European Conference on Computer Vision (ECCV 2016), Springer, 2016, pp. 852–

  98. [107]

    T. Chen, W. Yu, R. Chen, L. Lin, Knowledge-embedded rout- ing network for scene graph generation, in: IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2019), Computer Vision Foundation / IEEE, 2019, pp. 6163–6171. doi:10.1109/CVPR.2019.00632

  99. [108]

    Y . Hu, A. Chapman, G. Wen, D. W. Hall, What Can Knowl- edge Bring to Machine Learning?—A Survey of Low-shot Learning for Structured Data, ACM Trans. Intell. Syst. Tech- nol. 13 (2022) 1–45. doi:10.1145/3510030

  100. [109]

    W. Ren, Y . Tang, Q. Sun, C. Zhao, Q.-L. Han, Vi- sual semantic segmentation based on few /zero-shot learn- ing: An overview, IEEE /CAA J. Autom. Sin. (2023) 1–21. doi:10.1109/JAS.2023.123207

  101. [110]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, I. Sutskever, Learning transferable visual models from natural language supervision, in: Proceedings of the 38 th International Conference on Machine Learn...

  102. [111]

    S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, L. Zhang, Grounding DINO: mar- rying DINO with grounded pre-training for open-set object detection, CoRR abs /2303.05499 (2023). doi: 10.48550/ ARXIV.2303.05499. arXiv:2303.05499

  103. [112]

    T. Ren, Q. Jiang, S. Liu, Z. Zeng, W. Liu, H. Gao, H. Huang, Z. Ma, X. Jiang, Y . Chen, Y . Xiong, H. Zhang, F. Li, P. Tang, K. Yu, L. Zhang, Grounding DINO 1.5: Advance the ”edge” of open-set object detection, CoRR abs /2405.10300 (2024). doi:10.48550/ARXIV.2405.10300. arXiv:...

  104. [114]

    F. Zeng, W. Gan, Y . Wang, N. Liu, P. S. Yu, Large language models for robotics: A survey, CoRR abs /2311.07226 (2023). doi:10.48550/ARXIV.2311.07226. arXiv:2311.07226

  105. [115]

    D. Liu, R. Zhang, L. Qiu, S. Huang, W. Lin, S. Zhao, S. Geng, Z. Lin, P. Jin, K. Zhang, W. Shao, C. Xu, C. He, J. He, H. Shao, P. Lu, Y . Qiao, H. Li, P. Gao, SPHINX-X: scaling data and parameters for a family of multi-modal large language models, in: Forty-first International...

  106. [116]

    R. Anil, S. Borgeaud, Y . Wu, J. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, D. Sil- ver, S. Petrov, M. Johnson, I. Antonoglou, J. Schrittwieser, A. Glaese, J. Chen, E. Pitler, T. P. Lillicrap, A. Lazaridou, O. Firat, J. Molloy, M. Isard, P. R. ...

  107. [117]

    J. Li, D. Li, C. Xiong, S. C. H. Hoi, BLIP: bootstrapping language-image pre-training for unified vision-language un- derstanding and generation, in: International Conference on Machine Learning (ICML 2022), volume 162 of Proceed- ings of Machine Learning Research, PMLR, 2022,...

  108. [118]

    Z. Yang, L. Li, K. Lin, J. Wang, C. Lin, Z. Liu, L. Wang, The dawn of LMMs: Preliminary explorations with GPT- 4V(ision), CoRR abs /2309.17421 (2023). doi: 10.48550/ ARXIV.2309.17421. arXiv:2309.17421

  109. [119]

    F. Li, R. Zhang, H. Zhang, Y . Zhang, B. Li, W. Li, Z. Ma, C. Li, LLaV A-NeXT-Interleave: Tackling multi-image, video, and 3D in large multimodal models, CoRR abs/2407.07895 (2024). 21 doi:10.48550/ARXIV.2407.07895. arXiv:2407.07895

  110. [120]

    C. Jia, Y . Yang, Y . Xia, Y . Chen, Z. Parekh, H. Pham, Q. V . Le, Y . Sung, Z. Li, T. Duerig, Scaling up visual and vision- language representation learning with noisy text supervision, in: Proceedings of the 38 th International Conference on Ma- chine Learning (ICML 2021), ...

  111. [121]

    J. Li, D. Li, S. Savarese, S. C. H. Hoi, BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models, in: International Conference on Ma- chine Learning (ICML 2023), volume 202 of Proceedings of Machine Learning Research, PMLR, 2023, ...

  112. [123]

    Ghiasi, X

    G. Ghiasi, X. Gu, Y . Cui, T.-Y . Lin, Scaling open-vocabulary image segmentation with image-level labels, in: European Conference on Computer Vision, Springer, 2022, pp. 540–557

  113. [124]

    Z. Ding, J. Wang, Z. Tu, Open-vocabulary universal image segmentation with MaskCLIP, in: International Conference on Machine Learning (ICML 2023), volume 202 of Proceedings of Machine Learning Research, PMLR, 2023, pp. 8090–8102

  114. [125]

    X. Zou, Z. Dou, J. Yang, Z. Gan, L. Li, C. Li, X. Dai, H. Behl, J. Wang, L. Yuan, N. Peng, L. Wang, Y . J. Lee, J. Gao, Generalized decoding for pixel, image, and language, in: IEEE /CVF Conference on Computer Vision and Pattern Recognition (CVPR 2023), IEEE, 2023, pp. 15116–1...

  115. [126]

    B. Li, K. Q. Weinberger, S. J. Belongie, V . Koltun, R. Ran- ftl, Language-driven semantic segmentation, in: The Tenth International Conference on Learning Representations (ICLR 2022), 2022

  116. [127]

    M. Xu, Z. Zhang, F. Wei, H. Hu, X. Bai, SAN: side adapter network for open-vocabulary semantic segmentation, IEEE Trans. Pattern Anal. Mach. Intell. 45 (2023) 15546–15561. doi:10.1109/TPAMI.2023.3311618

  117. [130]

    J. Kerr, C. M. Kim, K. Goldberg, A. Kanazawa, M. Tan- cik, LERF: language embedded radiance fields, in: IEEE/CVF International Conference on Computer Vision (ICCV 2023), IEEE, 2023, pp. 19672–19682. doi: 10.1109/ ICCV51070.2023.01807

  118. [131]

    N. M. M. Shafiullah, C. Paxton, L. Pinto, S. Chintala, A. Szlam, CLIP-Fields: Weakly supervised semantic fields for robotic memory, in: Robotics: Science and Systems XIX (RSS 2023), 2023. doi:10.15607/RSS.2023.XIX.074

  119. [133]

    S. Peng, K. Genova, OpenScene: 3D Scene Understanding With Open V ocabularies, in: Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 815–824

  120. [134]

    K. M. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, G. Iyer, S. Saryazdi, T. Chen, A. Maalouf, S. Li, N. V . Keetha, A. Tewari, J. B. Tenenbaum, C. M. de Melo, K. M. Krishna, L. Paull, F. Shkurti, A. Torralba, ConceptFusion: Open-set multimodal 3D mapping, in: Robotics: Sci...

  121. [135]

    Cheng, I

    B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, R. Girdhar, Masked-attention mask transformer for universal image seg- mentation, in: IEEE /CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022), IEEE, 2022, pp. 1280–

  122. [136]

    P. Liu, Y . Orru, J. Vakil, C. Paxton, N. M. M. Shafiullah, L. Pinto, OK-Robot: What really matters in integrating open- knowledge models for robotics, in: 2 nd Workshop on Mobile Manipulation and Embodied Intelligence at ICRA 2024, 2024

  123. [137]

    X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Wang, L. Wang, J. Gao, Y . J. Lee, Segment everything everywhere all at once, Adv. Neural Inf. Process. Syst. 36 (2024)

  124. [138]

    Werby, C

    A. Werby, C. Huang, M. B ¨uchner, A. Valada, W. Burgard, Hierarchical open-vocabulary 3d scene graphs for language- grounded robot navigation, in: First Workshop on Vision- Language Models for Navigation and Manipulation at ICRA 2024, 2024

  125. [139]

    Q. Gu, A. Kuwajerwala, S. Morin, K. M. Jatavallab- hula, B. Sen, A. Agarwal, C. Rivera, W. Paul, K. Ellis, R. Chellappa, C. Gan, C. M. de Melo, J. B. Tenenbaum, A. Torralba, F. Shkurti, L. Paull, ConceptGraphs: Open- vocabulary 3D scene graphs for perception and planning, in: ...

  126. [140]

    Maggio, Y

    D. Maggio, Y . Chang, N. Hughes, M. Trang, J. D. Grif- fith, C. Dougherty, E. Cristofalo, L. Schmid, L. Carlone, Clio: Real-time task-driven open-set 3D scene graphs, IEEE Robotics Autom. Lett. 9 (2024) 8921–8928. doi: 10.1109/ LRA.2024.3451395

  127. [141]

    Tancik, E

    M. Tancik, E. Weber, E. Ng, R. Li, B. Yi, T. Wang, A. Kristof- fersen, J. Austin, K. Salahi, A. Ahuja, et al., Nerfstudio: A modular framework for neural radiance field development, in: ACM SIGGRAPH 2023 Conference Proceedings, 2023, pp. 1– 12

  128. [142]

    Medeiros, Language segment-anything, [software], 2024

    L. Medeiros, Language segment-anything, [software], 2024. URL: https://github.com/luca-medeiros/lang- segment-anything

  129. [143]

    Handa, T

    A. Handa, T. Whelan, J. McDonald, A. Davison, A benchmark for RGB-D visual odometry, 3D reconstruction and SLAM, in: IEEE Intl. Conf. on Robotics and Automation, ICRA, Hong Kong, China, 2014

  130. [144]

    doi: 10.48550/ARXIV.2303.08774

    OpenAI, GPT-4 technical report, CoRR abs/2303.08774 (2023). doi: 10.48550/ARXIV.2303.08774. arXiv:2303.08774

  131. [145]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, I. Polosukhin, Attention is all you need, in: Advances in Neural Information Processing Systems, volume 30, 2017

  132. [146]

    Kambhampati, Can large language models reason and plan?, Ann

    S. Kambhampati, Can large language models reason and plan?, Ann. N.Y . Acad. Sci. 1534 (2024) 15–18. doi:10.1111/ nyas.15125

  133. [147]

    K. Rana, J. Haviland, S. Garg, J. Abou-Chakra, I. D. Reid, N. S ¨underhauf, SayPlan: Grounding large language models using 3D scene graphs for scalable robot task planning, in: Conference on Robot Learning (CoRL 2023), volume 229 of Proceedings of Machine Learning Research, 20...

  134. [148]

    Ocker, J

    F. Ocker, J. Deigm ¨oller, J. Eggert, Exploring large lan- guage models as a source of common-sense knowledge 22 for robots, CoRR abs /2311.08412 (2023). doi: 10.48550/ ARXIV.2311.08412. arXiv:2311.08412

  135. [149]

    S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, E. Chen, A survey on multimodal large language models, CoRR abs/2306.13549 (2023). doi: 10.48550/ARXIV.2306.13549. arXiv:2306.13549

  136. [150]

    J. Wu, W. Gan, Z. Chen, S. Wan, S. Y . Philip, Multimodal large language models: A survey, in: 2023 IEEE International Con- ference on Big Data (BigData), IEEE, 2023, pp. 2247–2256

  137. [151]

    A. Z. Ren, J. Clark, A. Dixit, M. Itkina, A. Majumdar, D. Sadigh, Explore until confident: E fficient exploration for embodied question answering, in: Proceedings of Robotics: Science and Systems 2024, 2024. doi: 10.15607/ RSS.2024.XX.089

  138. [152]

    H. Liu, C. Li, Q. Wu, Y . J. Lee, Visual instruction tuning, in: A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, S. Levine (Eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Pro- cessing Systems (NeurIPS 2023), 2023

  139. [153]

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B....

  140. [154]

    Honerkamp, M

    D. Honerkamp, M. B ¨uchner, F. Despinoy, T. Welschehold, A. Valada, Language-grounded dynamic scene graphs for interactive object search with mobile manipulation, IEEE Robotics Autom. Lett. 9 (2024) 8298–8305. doi: 10.1109/ LRA.2024.3441495

  141. [155]

    Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y . Xu, E. Ishii, Y . Bang, A. Madotto, P. Fung, Survey of hallucination in natural lan- guage generation, ACM Comput. Surv. 55 (2023) 248:1– 248:38. doi:10.1145/3571730

  142. [156]

    Valmeekam, A

    K. Valmeekam, A. Olmo, S. Sreedharan, S. Kambhampati, Large language models still can’t plan (a benchmark for LLMs on planning and reasoning about change), in: NeurIPS 2022 Foundation Models for Decision Making Workshop, 2022

  143. [157]

    Rudin, Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead, Nat

    C. Rudin, Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead, Nat. Mach. Intell. 1 (2019) 206–215. doi: 10.1038/S42256- 019-0048-X

  144. [158]

    Huang, W

    L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, T. Liu, A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, CoRR abs/2311.05232 (2023). doi:10.48550/ARXIV.2311.05232. arXiv:2311.05232

  145. [159]

    Valmeekam, K

    K. Valmeekam, K. Stechly, S. Kambhampati, LLMs still can’t plan; can LRMs? A preliminary evaluation of OpenAI’s o1 on PlanBench, CoRR abs /2409.13373 (2024). doi:10.48550/ ARXIV.2409.13373. arXiv:2409.13373

  146. [160]

    Braunschweig, M

    B. Braunschweig, M. Ghallab (Eds.), Reflections on Artifi- cial Intelligence for Humanity, volume 12600 ofLecture Notes in Computer Science , Springer, 2021. doi: 10.1007/978-3- 030-69128-8

  147. [161]

    Helberger, N

    N. Helberger, N. Diakopoulos, ChatGPT and the AI Act, In- ternet Policy Rev. 12 (2023). doi:10.14763/2023.1.1682

  148. [162]

    J. Jiao, S. Afroogh, Y . Xu, C. Phillips, Navigating LLM ethics: Advancements, challenges, and future directions, CoRR abs/2406.18841 (2024). doi: 10.48550/ARXIV.2406.18841. arXiv:2406.18841. 23

  149. [206]

    doi:10.1007/S10514-012-9321-0

  150. [869]

    doi:10.1007/978-3-319-46448-0\ 51

  151. [1176]

    doi:10.1007/978-3-319-32552-1\ 46

  152. [1289]

    doi:10.1109/CVPR52688.2022.00135

  153. [2014]

    doi:10.15607/RSS.2014.X.007

  154. [4765]

    doi:10.1109/IROS.2018.8594299

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.