Pith. sign in

REVIEW 2 cited by

HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.12955 v2 pith:5NI6FKOC submitted 2025-03-17 cs.CV

classification cs.CV
keywords his-gpthis-qahumanunderstandinghuman-in-scenemodelsscenebenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). Given a human motion within a 3D scene, HIS-QA requires the agent to comprehend human states and behaviors, reason about its surrounding environment, and answer human-related questions within the scene. To support this new task, we present HIS-Bench, a multimodal benchmark that systematically evaluates HIS understanding across a broad spectrum, from basic perception to commonsense reasoning and planning. Our evaluation of various vision-language models on HIS-Bench reveals significant limitations in their ability to handle HIS-QA tasks. To this end, we propose HIS-GPT, the first foundation model for HIS understanding. HIS-GPT integrates 3D scene context and human motion dynamics into large language models while incorporating specialized mechanisms to capture human-scene interactions. Extensive experiments demonstrate that HIS-GPT sets a new state-of-the-art on HIS-QA tasks. We hope this work inspires future research on human behavior analysis in 3D scenes, advancing embodied AI and world models. The codes and data: https://github.com/ZJHTerry18/HumanInScene.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Grounding Intelligence in Movement

    cs.AI 2025-07 conditional novelty 4.0 of 10

    Movement should be treated as a first-class AI modeling modality, and a unified, biomechanically grounded movement foundation model built from aggregated data across species and sensors is the proposed path forward.

  2. Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

    cs.CV 2025-02 conditional novelty 4.0 of 10

    A structured survey of 3D Scene Question Answering that categorizes datasets, methods, and metrics and finds a common encoder-fusion-prediction pipeline across approaches.

Pith tools