Pith. sign in

REVIEW 7 cited by

ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.18600 v2 pith:MQPU7LPO submitted 2024-12-24 cs.CV cs.GR

ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation

classification cs.CV cs.GR
keywords human-sceneinteractionsscenesinteractionzerohsidataenvironmentsgeneration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Human-scene interaction (HSI) generation is crucial for applications in embodied AI, virtual reality, and robotics. Yet, existing methods cannot synthesize interactions in unseen environments such as in-the-wild scenes or reconstructed scenes, as they rely on paired 3D scenes and captured human motion data for training, which are unavailable for unseen environments. We present ZeroHSI, a novel approach that enables zero-shot 4D human-scene interaction synthesis, eliminating the need for training on any MoCap data. Our key insight is to distill human-scene interactions from state-of-the-art video generation models, which have been trained on vast amounts of natural human movements and interactions, and use differentiable rendering to reconstruct human-scene interactions. ZeroHSI can synthesize realistic human motions in both static scenes and environments with dynamic objects, without requiring any ground-truth motion data. We evaluate ZeroHSI on a curated dataset of different types of various indoor and outdoor scenes with different interaction prompts, demonstrating its ability to generate diverse and contextually appropriate human-scene interactions.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative Refinement

    cs.CV 2026-04 unverdicted novelty 7.0

    InfBaGel generates consistent human-object-scene interactions via dynamic perception during iterative refinement in a consistency model, bump-aware guidance to avoid collisions, and hybrid training that mixes synthesi...

  2. GenHSI: Controllable Generation of Human-Scene Interaction Videos

    cs.CV 2025-06 unverdicted novelty 7.0

    GenHSI is a training-free three-stage pipeline that turns a scene image, character image, and complex HSI prompt into long videos with plausible chained interactions by generating atomic actions, 3D keyframes via 2D i...

  3. VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification

    cs.CV 2025-12 unverdicted novelty 6.0

    VHOI densifies sparse trajectories into color-encoded HOI mask sequences and conditions a fine-tuned video diffusion model on them to produce controllable human-object interaction videos, including full navigation sequences.

  4. Three ways to share a QPU: Scheduling strategies for hybrid Quantum-HPC applications

    quant-ph 2026-04 unverdicted novelty 5.0

    Three scheduling strategies for hybrid quantum-HPC systems cut classical resource use by up to 64% or boost QPU utilization depending on workload balance, validated on real hardware.

  5. Three ways to share a QPU: Scheduling strategies for hybrid Quantum-HPC applications

    quant-ph 2026-04 unverdicted novelty 5.0

    Three complementary HPC-QC scheduling strategies cut classical resource use by up to 64% or improve QPU utilization depending on quantum-classical workload balance.

  6. Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation

    cs.CV 2026-04 unverdicted novelty 5.0

    Prompt-driven image-to-video generation produces deictic gestures that match real data visually, add useful variety, and improve downstream recognition models when mixed with human recordings.

  7. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.