REVIEW 47 references
Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots
T0 review · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Social 3D Scene Graphs extend 3D scene graphs with human activities and relations, but the evaluation ground truth is derived from the model's own outputs.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
The method, called ReaSoN, uses a vision-language model to describe each person's behavior from an image, estimates head pose to determine where each person is looking, and uses a large language model to connect people to objects in the 3D map that match the described activity. It also introduces SocialGraph3D, a benchmark of 8 synthetic home scenes with 42 humans and over 100 annotated relationships.
The problem is that the ground truth relationships in the benchmark were not annotated independently. Instead, the authors let the same models being evaluated propose candidate activities, then had human raters judge which candidates were plausible. The final ground truth is a filtered union of these model-generated candidates. This means the top-scoring model, ReaSoN, is being measured against a list that largely consists of its own suggestions, which inflates its recall. The paper's main claim that its representation improves human activity prediction therefore rests on a circular evaluation.
Extended reading notes
Core claim
From the contributions: 'We use our benchmark to demonstrate that our representation outperform state-of-the-art performance on the introduced tasks.' This means ReaSoN, by augmenting 3D Scene Graphs with human activities and relationships, improves human activity prediction and social scene understanding compared to prior 3DSG methods. If correct, the representation enables socially aware behavior in robots.
Load-bearing premise
The SocialGraph3D ground truth for activities, constructed by filtering the models' own predictions through human plausibility ratings, is an accurate and unbiased measure of actual human activities. This enters in Sec. IV-B: 'The final GT is thus the union of these filtered predictions and consistently suggested activities.' If the GT is biased toward one model (e.g., because ReaSoN's predictions were all also suggested by evaluators, as the text notes), the reported F1 comparison in Table II is invalid, so the central claim of improved activity prediction falls.
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
free parameters (4)
- Pruning relative threshold τ =
0.4
- Pruning absolute cutoff N_min =
2
- GT plausibility threshold (median split) =
4.33
- Open-ended suggestion inclusion threshold =
30%
assumptions (5)
- domain assumption Humans look at the objects they engage with
- domain assumption Humans are static except facial/mouth movements
- domain assumption VLM and LLM (GPT-5) provide reliable commonsense reasoning
- domain assumption Synthetic scenes are representative of real environments
- domain assumption The human-in-the-loop GT is an accurate measure of actual activities
Cite this review
Pith. "Pith review of Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots." pith.science (2026). https://pith.science/paper/5HPFMMX4
@misc{pith2026250924966,
author = {Pith},
title = {Pith review of: Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots},
year = {2026},
howpublished = {\url{https://pith.science/paper/5HPFMMX4}},
note = {Machine review of arXiv:2509.24966}
}
read the original abstract
Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-aware ways. While 3D Scene Graphs have emerged as a powerful semantic representation for scene understanding, existing approaches largely ignore humans in the scene, also due to the lack of annotated human-environment relationships. Moreover, existing methods typically capture only open-vocabulary relations from single image frames, which limits their ability to model long-range interactions beyond the observed content. We introduce Social 3D Scene Graphs, an augmented 3D Scene Graph representation that captures humans, their attributes, activities and relationships in the environment, both local and remote, using an open-vocabulary framework. Furthermore, we introduce a new benchmark consisting of synthetic environments with comprehensive human-scene relationship annotations and diverse types of queries for evaluating social scene understanding in 3D. The experiments demonstrate that our representation improves human activity prediction and reasoning about human-environment relations, paving the way toward socially intelligent robots.
Figures
Reference graph
Works this paper leans on
-
[1]
Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age,
C. Cadena, L. Carlone, H. Carrillo, Y . Latif, D. Scaramuzza, J. Neira, I. Reid, and J. J. Leonard, “Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age,”IEEE Trans. on Robotics, 2016
2016
-
[2]
3d scene graph: A structure for unified semantics, 3d space, and camera,
I. Armeni, Z.-Y . He, J. Gwak, A. R. Zamir, M. Fischer, J. Malik, and S. Savarese, “3d scene graph: A structure for unified semantics, 3d space, and camera,”ICCV, 2019
2019
-
[3]
3-d scene graph: A sparse and semantic representation of physical environments for intelligent agents,
U.-H. Kim, J.-M. Park, T.-j. Song, and J.-H. Kim, “3-d scene graph: A sparse and semantic representation of physical environments for intelligent agents,”IEEE Trans. on Cybernetics, 2020
2020
-
[4]
Foundations of spatial perception for robotics: Hierarchical representations and real-time systems,
N. Hughes, Y . Chang, S. Hu, R. Talak, R. Abdulhai, J. Strader, and L. Carlone, “Foundations of spatial perception for robotics: Hierarchical representations and real-time systems,”IJRR, 2024
2024
-
[5]
Scannet: Richly-annotated 3d reconstructions of indoor scenes,
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” inCVPR, 2017
2017
-
[6]
The Replica dataset: A digital replica of indoor spaces,
J. Straub, T. Whelan, L. Ma, Y . Chen, and et al., “The Replica dataset: A digital replica of indoor spaces,”arXiv, 2019
2019
-
[7]
Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data,
G. Baruch, Z. Chen, A. Dehghan, T. Dimry, Y . Feigin, P. Fu, T. Gebauer, B. Joffe, D. Kurz, A. Schwartzet al., “Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data,”arXiv preprint arXiv:2111.08897, 2021
arXiv 2021
-
[8]
Scannet++: A high-fidelity dataset of 3d indoor scenes,
C. Yeshwanth, Y .-C. Liu, M. Nießner, and A. Dai, “Scannet++: A high-fidelity dataset of 3d indoor scenes,” inICCV, 2023
2023
Show all 47 references
-
[9]
Open-V ocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces,
C. Zhang, A. Delitzas, F. Wang, R. Zhang, X. Ji, M. Pollefeys, and F. Engelmann, “Open-V ocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces,” inCVPR, 2025
2025
-
[10]
Kimera: From slam to spatial perception with 3d dynamic scene graphs,
A. Rosinol, A. Violette, M. Abate, N. Hughes, Y . Chang, J. Shi, A. Gupta, and L. Carlone, “Kimera: From slam to spatial perception with 3d dynamic scene graphs,”IJRR, 2021
2021
-
[11]
Hydra: A real-time spatial perception system for 3d scene graph construction and optimization,
N. Hughes, Y . Chang, and L. Carlone, “Hydra: A real-time spatial perception system for 3d scene graph construction and optimization,” RSS, 2022
2022
-
[12]
Clio: Real-time task-driven open-set 3d scene graphs,
D. Maggio, Y . Chang, N. Hughes, M. Trang, D. Griffith, C. Dougherty, E. Cristofalo, L. Schmid, and L. Carlone, “Clio: Real-time task-driven open-set 3d scene graphs,”RA-L, 2024
2024
-
[13]
Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,
Q. Gu, A. Kuwajerwala, S. Morin, K. M. Jatavallabhula, B. Sen, A. Agarwal, C. Rivera, W. Paul, K. Ellis, R. Chellappa, C. Gan, C. M. de Melo, J. B. Tenenbaum, A. Torralba, F. Shkurti, and L. Paull, “Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,” i...
2024
-
[14]
Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,
A. Werby, C. Huang, M. Büchner, A. Valada, and W. Burgard, “Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,” inRSS, 2024
2024
-
[15]
Fungraph: Functionality aware 3d scene graphs for language-prompted scene interaction,
D. Rotondi, F. Scaparro, H. Blum, and K. O. Arras, “Fungraph: Functionality aware 3d scene graphs for language-prompted scene interaction,”IROS, 2025
2025
-
[16]
Learning 3D Semantic Scene Graphs from 3D Indoor Reconstructions,
J. Wald, H. Dhamo, N. Navab, and F. Tombari, “Learning 3D Semantic Scene Graphs from 3D Indoor Reconstructions,” inCVPR, 2020
2020
-
[17]
Scenegraph- fusion: Incremental 3d scene graph prediction from rgb-d sequences,
S.-C. Wu, J. Wald, K. Tateno, N. Navab, and F. Tombari, “Scenegraph- fusion: Incremental 3d scene graph prediction from rgb-d sequences,” inCVPR, 2021
2021
-
[18]
Knowledge-inspired 3d scene graph prediction in point cloud,
S. Zhang, A. Hao, H. Qinet al., “Knowledge-inspired 3d scene graph prediction in point cloud,”NeurIPS, 2021
2021
-
[19]
Vl- sat: Visual-linguistic semantics assisted training for 3d semantic scene graph prediction in point cloud,
Z. Wang, B. Cheng, L. Zhao, D. Xu, Y . Tang, and L. Sheng, “Vl- sat: Visual-linguistic semantics assisted training for 3d semantic scene graph prediction in point cloud,” inCVPR, 2023
2023
-
[20]
Open3dsg: Open-vocabulary 3d scene graphs from point clouds with queryable objects and open-set relationships,
S. Koch, N. Vaskevicius, M. Colosi, P. Hermosilla, and T. Ropinski, “Open3dsg: Open-vocabulary 3d scene graphs from point clouds with queryable objects and open-set relationships,” inCVPR, 2024
2024
-
[21]
Scene graph generation by iterative message passing,
D. Xu, Y . Zhu, C. B. Choy, and L. Fei-Fei, “Scene graph generation by iterative message passing,”CVPR, 2017
2017
-
[22]
On the difficulty of training recurrent neural networks,
R. Pascanu, T. Mikolov, and Y . Bengio, “On the difficulty of training recurrent neural networks,” inICML, 2013
2013
-
[23]
3d dynamic scene graphs: Actionable spatial perception with places, objects, and humans,
A. Rosinol, A. Gupta, M. Abate, J. Shi, and L. Carlone, “3d dynamic scene graphs: Actionable spatial perception with places, objects, and humans,”RSS, 2020
2020
-
[24]
Hierarchical representations and explicit memory: Learning effective navigation policies on 3d scene graphs using graph neural networks,
Z. Ravichandran, L. Peng, N. Hughes, J. D. Griffith, and L. Carlone, “Hierarchical representations and explicit memory: Learning effective navigation policies on 3d scene graphs using graph neural networks,” ICRA, 2022
2022
-
[25]
Long-term human trajectory prediction using 3d dynamic scene graphs,
N. Gorlo, L. Schmid, and L. Carlone, “Long-term human trajectory prediction using 3d dynamic scene graphs,” 2024
2024
-
[26]
Taskography: Evaluating robot task planning over large 3d scene graphs,
C. Agia, K. M. Jatavallabhula, M. Khodeir, O. Miksik, V . Vineet, M. Mukadam, L. Paull, and F. Shkurti, “Taskography: Evaluating robot task planning over large 3d scene graphs,”CoRL, 2022
2022
-
[27]
Anticipatory planning for performant long-lived robot in large-scale home-like environments,
M. R. H. Talukder, R. I. Arnob, and G. J. Stein, “Anticipatory planning for performant long-lived robot in large-scale home-like environments,”ICRA, 2025
2025
-
[28]
Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning,
K. Rana, J. Haviland, S. Garg, J. Abou-Chakra, I. D. Reid, and N. Sünderhauf, “Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning,”CoRL, 2023
2023
-
[29]
Delta: Decomposed efficient long-term robot task planning using large lan- guage models,
Y . Liu, L. Palmieri, S. Koch, I. Georgievski, and M. Aiello, “Delta: Decomposed efficient long-term robot task planning using large lan- guage models,”ICRA, 2025
2025
-
[30]
Saynav: Grounding large language models for dynamic planning to navigation in new environments,
A. Rajvanshi, K. Sikka, X. Lin, B. Lee, H.-P. Chiu, and A. Velasquez, “Saynav: Grounding large language models for dynamic planning to navigation in new environments,”AAAI, 2024
2024
-
[31]
Sg-nav: Online 3d scene graph prompting for llm-based zero-shot object navigation,
H. Yin, X. Xu, Z. Wu, J. Zhou, and J. Lu, “Sg-nav: Online 3d scene graph prompting for llm-based zero-shot object navigation,”NeurIPS, 2024
2024
-
[32]
Core challenges of social robot navigation: A survey,
C. Mavrogiannis, F. Baldini, A. Wang, D. Zhao, P. Trautman, A. Stein- feld, and J. Oh, “Core challenges of social robot navigation: A survey,” ACM Trans. on Human-Robot Interaction, 2023
2023
-
[33]
Learning socially normative robot naviga- tion behaviors with bayesian inverse reinforcement learning,
B. Okal and K. O. Arras, “Learning socially normative robot naviga- tion behaviors with bayesian inverse reinforcement learning,” inICRA, 2016
2016
-
[34]
A human aware mobile robot motion planner,
E. A. Sisbot, L. F. Marin-Urias, R. Alami, and T. Simeon, “A human aware mobile robot motion planner,”IEEE Trans. on Robotics, 2007
2007
-
[35]
Proactive model predictive control with multi-modal human motion prediction in cluttered dynamic environments,
L. Heuer, L. Palmieri, A. Rudenko, A. Mannucci, M. Magnusson, and K. O. Arras, “Proactive model predictive control with multi-modal human motion prediction in cluttered dynamic environments,” inIROS, 2023
2023
-
[36]
Moka: Open-world robotic manipulation through mark-based visual prompting,
K. Fang, F. Liu, P. Abbeel, and S. Levine, “Moka: Open-world robotic manipulation through mark-based visual prompting,”RSS, 2024
2024
-
[37]
Detection and quantification of occlusion for 3d human pose and shape estimation,
E. Girgin, B. Gökberk, and L. Akarun, “Detection and quantification of occlusion for 3d human pose and shape estimation,” inPattern Recognition. ICPR 2024 International Workshops and Challenges, S. Palaiahnakote, S. Schuckers, J.-M. Ogier, P. Bhattacharya, U. Pal, and S. Bhatt...
2024
-
[38]
Spatialrgpt: Grounded spatial reasoning in vision-language models,
A.-C. Cheng, H. Yin, Y . Fu, Q. Guo, R. Yang, J. Kautz, X. Wang, and S. Liu, “Spatialrgpt: Grounded spatial reasoning in vision-language models,”NeurIPS, 2025
2025
-
[39]
Verbatlas: a novel large-scale verbal semantic resource and its application to semantic role labeling,
A. Di Fabio, S. Conia, and R. Navigli, “Verbatlas: a novel large-scale verbal semantic resource and its application to semantic role labeling,” inEMNLP-IJCNLP, 2019
2019
-
[40]
Nounatlas: Filling the gap in nominal semantic role la- beling,
R. Navigli, M. Pinto, P. Silvestri, D. Rotondi, S. Ciciliano, and A. Scirè, “Nounatlas: Filling the gap in nominal semantic role la- beling,” inACL, 2024
2024
-
[41]
Yolo-world: Real-time open-vocabulary object detection,
T. Cheng, L. Song, Y . Ge, W. Liu, X. Wang, and Y . Shan, “Yolo-world: Real-time open-vocabulary object detection,”CVPR, 2024
2024
-
[42]
Segment anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollár, and R. Gir- shick, “Segment anything,” inICCV, 2023
2023
-
[43]
Mobilenets: Efficient convolutional neural networks for mobile vision applications,
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” 2017
2017
-
[44]
Toward robust and unconstrained full range of rotation head pose estimation,
T. Hempel, A. A. Abdelrahman, and A. Al-Hamadi, “Toward robust and unconstrained full range of rotation head pose estimation,”IEEE Trans. on Image Processing, 2024
2024
- [45]
-
[46]
Learning transferable visual models from natural language supervi- sion,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervi- sion,” inProceedings of the 38th Int. Conf. on Machine Learning, 2021
2021
-
[47]
Ufomap: An efficient probabilistic 3d mapping framework that embraces the unknown,
D. Duberg and P. Jensfelt, “Ufomap: An efficient probabilistic 3d mapping framework that embraces the unknown,”IEEE Robotics and Automation Letters, 2020
2020
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.