Pith. sign in

REVIEW 47 references

Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots

T0 review · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read Social 3D Scene Graphs extend 3D scene graphs with human activities and relations, but the evaluation ground truth is derived from the model's own outputs.

arxiv 2509.24966 v2 pith:5HPFMMX4 submitted 2025-09-29 cs.CV

classification cs.CV
keywords scenegraphsrelationsrepresentationrobotssocialunderstandingexisting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Service robots need to understand not just where objects are, but what people are doing and how they relate to objects and each other. 3D Scene Graphs are a structured map of a room that lists objects, rooms, and their spatial relationships. This paper adds social information: human nodes, their activities such as sitting, talking, or watching TV, and the objects or people they interact with, even when the object is outside the camera's view.

The method, called ReaSoN, uses a vision-language model to describe each person's behavior from an image, estimates head pose to determine where each person is looking, and uses a large language model to connect people to objects in the 3D map that match the described activity. It also introduces SocialGraph3D, a benchmark of 8 synthetic home scenes with 42 humans and over 100 annotated relationships.

The problem is that the ground truth relationships in the benchmark were not annotated independently. Instead, the authors let the same models being evaluated propose candidate activities, then had human raters judge which candidates were plausible. The final ground truth is a filtered union of these model-generated candidates. This means the top-scoring model, ReaSoN, is being measured against a list that largely consists of its own suggestions, which inflates its recall. The paper's main claim that its representation improves human activity prediction therefore rests on a circular evaluation.

Extended reading notes

Core claim

From the contributions: 'We use our benchmark to demonstrate that our representation outperform state-of-the-art performance on the introduced tasks.' This means ReaSoN, by augmenting 3D Scene Graphs with human activities and relationships, improves human activity prediction and social scene understanding compared to prior 3DSG methods. If correct, the representation enables socially aware behavior in robots.

Load-bearing premise

The SocialGraph3D ground truth for activities, constructed by filtering the models' own predictions through human plausibility ratings, is an accurate and unbiased measure of actual human activities. This enters in Sec. IV-B: 'The final GT is thus the union of these filtered predictions and consistently suggested activities.' If the GT is biased toward one model (e.g., because ReaSoN's predictions were all also suggested by evaluators, as the text notes), the reported F1 comparison in Table II is invalid, so the central claim of improved activity prediction falls.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The paper relies on several domain assumptions about human behavior and tool reliability. Most critically, the ground-truth construction assumes that plausibility ratings of the models' own outputs can serve as ground truth, which is the source of the circular evaluation.

free parameters (4)
  • Pruning relative threshold τ = 0.4
    Used in Eq. 1 to prune spurious activity edges; set without systematic tuning, directly affects precision/recall.
  • Pruning absolute cutoff N_min = 2
    Minimum number of detections to retain an edge, Eq. 1; hand-set, affects final graph.
  • GT plausibility threshold (median split) = 4.33
    Activities with average human rating above 4.33 (median of ratings) are kept in the GT; this threshold determines which model predictions enter the ground truth.
  • Open-ended suggestion inclusion threshold = 30%
    An activity suggested by evaluators is added to GT if at least 30% of evaluators proposed it; hand-set, influences GT composition.
assumptions (5)
  • domain assumption Humans look at the objects they engage with
    Stated in the Introduction: 'by assuming that humans look at the objects they engage with'. Underpins the interaction context estimator and activity solver that use head pose to link remote activities.
  • domain assumption Humans are static except facial/mouth movements
    Stated in Sec. III-A: 'we assume that in these images, the humans remain still except for facial expressions and mouth movements'. Restricts the benchmark and method to near-static scenes.
  • domain assumption VLM and LLM (GPT-5) provide reliable commonsense reasoning
    The Activity Descriptor and Solver depend on proprietary GPT-5 responses; their reliability is assumed, not verified in the paper.
  • domain assumption Synthetic scenes are representative of real environments
    The benchmark uses 8 Unity synthetic homes; the paper claims generalization to real scenes based solely on the fact that pretrained models were trained on real data, not on empirical evidence.
  • domain assumption The human-in-the-loop GT is an accurate measure of actual activities
    Sec IV-B assumes plausibility ratings of model-generated candidates constitute ground truth. This is the assumption that creates circularity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots." pith.science (2026). https://pith.science/paper/5HPFMMX4

@misc{pith2026250924966,
  author       = {Pith},
  title        = {Pith review of: Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5HPFMMX4}},
  note         = {Machine review of arXiv:2509.24966}
}
read the original abstract

Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-aware ways. While 3D Scene Graphs have emerged as a powerful semantic representation for scene understanding, existing approaches largely ignore humans in the scene, also due to the lack of annotated human-environment relationships. Moreover, existing methods typically capture only open-vocabulary relations from single image frames, which limits their ability to model long-range interactions beyond the observed content. We introduce Social 3D Scene Graphs, an augmented 3D Scene Graph representation that captures humans, their attributes, activities and relationships in the environment, both local and remote, using an open-vocabulary framework. Furthermore, we introduce a new benchmark consisting of synthetic environments with comprehensive human-scene relationship annotations and diverse types of queries for evaluating social scene understanding in 3D. The experiments demonstrate that our representation improves human activity prediction and reasoning about human-environment relations, paving the way toward socially intelligent robots.

Figures

Figures reproduced from arXiv: 2509.24966 by the authors.

Figure 1
Figure 1. Example of a Social 3D Scene Graph. Each entity is [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. An overview of our solution ReaSoN, which extends existing 3D Scene Graphs with humans and their activities [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Illustrative example from our Social Scene Under [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Socially aware planning in a real-world scene based [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 2 linked inside Pith

  1. [1]

    Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age,

    C. Cadena, L. Carlone, H. Carrillo, Y . Latif, D. Scaramuzza, J. Neira, I. Reid, and J. J. Leonard, “Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age,”IEEE Trans. on Robotics, 2016

  2. [2]

    3d scene graph: A structure for unified semantics, 3d space, and camera,

    I. Armeni, Z.-Y . He, J. Gwak, A. R. Zamir, M. Fischer, J. Malik, and S. Savarese, “3d scene graph: A structure for unified semantics, 3d space, and camera,”ICCV, 2019

  3. [3]

    3-d scene graph: A sparse and semantic representation of physical environments for intelligent agents,

    U.-H. Kim, J.-M. Park, T.-j. Song, and J.-H. Kim, “3-d scene graph: A sparse and semantic representation of physical environments for intelligent agents,”IEEE Trans. on Cybernetics, 2020

  4. [4]

    Foundations of spatial perception for robotics: Hierarchical representations and real-time systems,

    N. Hughes, Y . Chang, S. Hu, R. Talak, R. Abdulhai, J. Strader, and L. Carlone, “Foundations of spatial perception for robotics: Hierarchical representations and real-time systems,”IJRR, 2024

  5. [5]

    Scannet: Richly-annotated 3d reconstructions of indoor scenes,

    A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” inCVPR, 2017

  6. [6]

    The Replica dataset: A digital replica of indoor spaces,

    J. Straub, T. Whelan, L. Ma, Y . Chen, and et al., “The Replica dataset: A digital replica of indoor spaces,”arXiv, 2019

  7. [7]

    Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data,

    G. Baruch, Z. Chen, A. Dehghan, T. Dimry, Y . Feigin, P. Fu, T. Gebauer, B. Joffe, D. Kurz, A. Schwartzet al., “Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data,”arXiv preprint arXiv:2111.08897, 2021

  8. [8]

    Scannet++: A high-fidelity dataset of 3d indoor scenes,

    C. Yeshwanth, Y .-C. Liu, M. Nießner, and A. Dai, “Scannet++: A high-fidelity dataset of 3d indoor scenes,” inICCV, 2023

Show all 47 references
  1. [9]

    Open-V ocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces,

    C. Zhang, A. Delitzas, F. Wang, R. Zhang, X. Ji, M. Pollefeys, and F. Engelmann, “Open-V ocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces,” inCVPR, 2025

  2. [10]

    Kimera: From slam to spatial perception with 3d dynamic scene graphs,

    A. Rosinol, A. Violette, M. Abate, N. Hughes, Y . Chang, J. Shi, A. Gupta, and L. Carlone, “Kimera: From slam to spatial perception with 3d dynamic scene graphs,”IJRR, 2021

  3. [11]

    Hydra: A real-time spatial perception system for 3d scene graph construction and optimization,

    N. Hughes, Y . Chang, and L. Carlone, “Hydra: A real-time spatial perception system for 3d scene graph construction and optimization,” RSS, 2022

  4. [12]

    Clio: Real-time task-driven open-set 3d scene graphs,

    D. Maggio, Y . Chang, N. Hughes, M. Trang, D. Griffith, C. Dougherty, E. Cristofalo, L. Schmid, and L. Carlone, “Clio: Real-time task-driven open-set 3d scene graphs,”RA-L, 2024

  5. [13]

    Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,

    Q. Gu, A. Kuwajerwala, S. Morin, K. M. Jatavallabhula, B. Sen, A. Agarwal, C. Rivera, W. Paul, K. Ellis, R. Chellappa, C. Gan, C. M. de Melo, J. B. Tenenbaum, A. Torralba, F. Shkurti, and L. Paull, “Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,” i...

  6. [14]

    Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,

    A. Werby, C. Huang, M. Büchner, A. Valada, and W. Burgard, “Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,” inRSS, 2024

  7. [15]

    Fungraph: Functionality aware 3d scene graphs for language-prompted scene interaction,

    D. Rotondi, F. Scaparro, H. Blum, and K. O. Arras, “Fungraph: Functionality aware 3d scene graphs for language-prompted scene interaction,”IROS, 2025

  8. [16]

    Learning 3D Semantic Scene Graphs from 3D Indoor Reconstructions,

    J. Wald, H. Dhamo, N. Navab, and F. Tombari, “Learning 3D Semantic Scene Graphs from 3D Indoor Reconstructions,” inCVPR, 2020

  9. [17]

    Scenegraph- fusion: Incremental 3d scene graph prediction from rgb-d sequences,

    S.-C. Wu, J. Wald, K. Tateno, N. Navab, and F. Tombari, “Scenegraph- fusion: Incremental 3d scene graph prediction from rgb-d sequences,” inCVPR, 2021

  10. [18]

    Knowledge-inspired 3d scene graph prediction in point cloud,

    S. Zhang, A. Hao, H. Qinet al., “Knowledge-inspired 3d scene graph prediction in point cloud,”NeurIPS, 2021

  11. [19]

    Vl- sat: Visual-linguistic semantics assisted training for 3d semantic scene graph prediction in point cloud,

    Z. Wang, B. Cheng, L. Zhao, D. Xu, Y . Tang, and L. Sheng, “Vl- sat: Visual-linguistic semantics assisted training for 3d semantic scene graph prediction in point cloud,” inCVPR, 2023

  12. [20]

    Open3dsg: Open-vocabulary 3d scene graphs from point clouds with queryable objects and open-set relationships,

    S. Koch, N. Vaskevicius, M. Colosi, P. Hermosilla, and T. Ropinski, “Open3dsg: Open-vocabulary 3d scene graphs from point clouds with queryable objects and open-set relationships,” inCVPR, 2024

  13. [21]

    Scene graph generation by iterative message passing,

    D. Xu, Y . Zhu, C. B. Choy, and L. Fei-Fei, “Scene graph generation by iterative message passing,”CVPR, 2017

  14. [22]

    On the difficulty of training recurrent neural networks,

    R. Pascanu, T. Mikolov, and Y . Bengio, “On the difficulty of training recurrent neural networks,” inICML, 2013

  15. [23]

    3d dynamic scene graphs: Actionable spatial perception with places, objects, and humans,

    A. Rosinol, A. Gupta, M. Abate, J. Shi, and L. Carlone, “3d dynamic scene graphs: Actionable spatial perception with places, objects, and humans,”RSS, 2020

  16. [24]

    Hierarchical representations and explicit memory: Learning effective navigation policies on 3d scene graphs using graph neural networks,

    Z. Ravichandran, L. Peng, N. Hughes, J. D. Griffith, and L. Carlone, “Hierarchical representations and explicit memory: Learning effective navigation policies on 3d scene graphs using graph neural networks,” ICRA, 2022

  17. [25]

    Long-term human trajectory prediction using 3d dynamic scene graphs,

    N. Gorlo, L. Schmid, and L. Carlone, “Long-term human trajectory prediction using 3d dynamic scene graphs,” 2024

  18. [26]

    Taskography: Evaluating robot task planning over large 3d scene graphs,

    C. Agia, K. M. Jatavallabhula, M. Khodeir, O. Miksik, V . Vineet, M. Mukadam, L. Paull, and F. Shkurti, “Taskography: Evaluating robot task planning over large 3d scene graphs,”CoRL, 2022

  19. [27]

    Anticipatory planning for performant long-lived robot in large-scale home-like environments,

    M. R. H. Talukder, R. I. Arnob, and G. J. Stein, “Anticipatory planning for performant long-lived robot in large-scale home-like environments,”ICRA, 2025

  20. [28]

    Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning,

    K. Rana, J. Haviland, S. Garg, J. Abou-Chakra, I. D. Reid, and N. Sünderhauf, “Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning,”CoRL, 2023

  21. [29]

    Delta: Decomposed efficient long-term robot task planning using large lan- guage models,

    Y . Liu, L. Palmieri, S. Koch, I. Georgievski, and M. Aiello, “Delta: Decomposed efficient long-term robot task planning using large lan- guage models,”ICRA, 2025

  22. [30]

    Saynav: Grounding large language models for dynamic planning to navigation in new environments,

    A. Rajvanshi, K. Sikka, X. Lin, B. Lee, H.-P. Chiu, and A. Velasquez, “Saynav: Grounding large language models for dynamic planning to navigation in new environments,”AAAI, 2024

  23. [31]

    Sg-nav: Online 3d scene graph prompting for llm-based zero-shot object navigation,

    H. Yin, X. Xu, Z. Wu, J. Zhou, and J. Lu, “Sg-nav: Online 3d scene graph prompting for llm-based zero-shot object navigation,”NeurIPS, 2024

  24. [32]

    Core challenges of social robot navigation: A survey,

    C. Mavrogiannis, F. Baldini, A. Wang, D. Zhao, P. Trautman, A. Stein- feld, and J. Oh, “Core challenges of social robot navigation: A survey,” ACM Trans. on Human-Robot Interaction, 2023

  25. [33]

    Learning socially normative robot naviga- tion behaviors with bayesian inverse reinforcement learning,

    B. Okal and K. O. Arras, “Learning socially normative robot naviga- tion behaviors with bayesian inverse reinforcement learning,” inICRA, 2016

  26. [34]

    A human aware mobile robot motion planner,

    E. A. Sisbot, L. F. Marin-Urias, R. Alami, and T. Simeon, “A human aware mobile robot motion planner,”IEEE Trans. on Robotics, 2007

  27. [35]

    Proactive model predictive control with multi-modal human motion prediction in cluttered dynamic environments,

    L. Heuer, L. Palmieri, A. Rudenko, A. Mannucci, M. Magnusson, and K. O. Arras, “Proactive model predictive control with multi-modal human motion prediction in cluttered dynamic environments,” inIROS, 2023

  28. [36]

    Moka: Open-world robotic manipulation through mark-based visual prompting,

    K. Fang, F. Liu, P. Abbeel, and S. Levine, “Moka: Open-world robotic manipulation through mark-based visual prompting,”RSS, 2024

  29. [37]

    Detection and quantification of occlusion for 3d human pose and shape estimation,

    E. Girgin, B. Gökberk, and L. Akarun, “Detection and quantification of occlusion for 3d human pose and shape estimation,” inPattern Recognition. ICPR 2024 International Workshops and Challenges, S. Palaiahnakote, S. Schuckers, J.-M. Ogier, P. Bhattacharya, U. Pal, and S. Bhatt...

  30. [38]

    Spatialrgpt: Grounded spatial reasoning in vision-language models,

    A.-C. Cheng, H. Yin, Y . Fu, Q. Guo, R. Yang, J. Kautz, X. Wang, and S. Liu, “Spatialrgpt: Grounded spatial reasoning in vision-language models,”NeurIPS, 2025

  31. [39]

    Verbatlas: a novel large-scale verbal semantic resource and its application to semantic role labeling,

    A. Di Fabio, S. Conia, and R. Navigli, “Verbatlas: a novel large-scale verbal semantic resource and its application to semantic role labeling,” inEMNLP-IJCNLP, 2019

  32. [40]

    Nounatlas: Filling the gap in nominal semantic role la- beling,

    R. Navigli, M. Pinto, P. Silvestri, D. Rotondi, S. Ciciliano, and A. Scirè, “Nounatlas: Filling the gap in nominal semantic role la- beling,” inACL, 2024

  33. [41]

    Yolo-world: Real-time open-vocabulary object detection,

    T. Cheng, L. Song, Y . Ge, W. Liu, X. Wang, and Y . Shan, “Yolo-world: Real-time open-vocabulary object detection,”CVPR, 2024

  34. [42]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollár, and R. Gir- shick, “Segment anything,” inICCV, 2023

  35. [43]

    Mobilenets: Efficient convolutional neural networks for mobile vision applications,

    A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” 2017

  36. [44]

    Toward robust and unconstrained full range of rotation head pose estimation,

    T. Hempel, A. A. Abdelrahman, and A. Al-Hamadi, “Toward robust and unconstrained full range of rotation head pose estimation,”IEEE Trans. on Image Processing, 2024

  37. [45]

    GPT-4 technical report,

    OpenAI, “GPT-4 technical report,”CoRR, vol. abs/2303.08774, 2023

  38. [46]

    Learning transferable visual models from natural language supervi- sion,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervi- sion,” inProceedings of the 38th Int. Conf. on Machine Learning, 2021

  39. [47]

    Ufomap: An efficient probabilistic 3d mapping framework that embraces the unknown,

    D. Duberg and P. Jensfelt, “Ufomap: An efficient probabilistic 3d mapping framework that embraces the unknown,”IEEE Robotics and Automation Letters, 2020

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.