Pith. sign in

hub

Hunyuanworld 1.0: Generating immersive, explorable, and interactive 3d worlds from words or pixels

19 Pith papers cite this work. Polarity classification is still indexing.

19 Pith papers citing it
abstract

Creating immersive and playable 3D worlds from texts or images remains a fundamental challenge in computer vision and graphics. Existing world generation approaches typically fall into two categories: video-based methods that offer rich diversity but lack 3D consistency and rendering efficiency, and 3D-based methods that provide geometric consistency but struggle with limited training data and memory-inefficient representations. To address these limitations, we present HunyuanWorld 1.0, a novel framework that combines the best of both worlds for generating immersive, explorable, and interactive 3D scenes from text and image conditions. Our approach features three key advantages: 1) 360{\deg} immersive experiences via panoramic world proxies; 2) mesh export capabilities for seamless compatibility with existing computer graphics pipelines; 3) disentangled object representations for augmented interactivity. The core of our framework is a semantically layered 3D mesh representation that leverages panoramic images as 360{\deg} world proxies for semantic-aware world decomposition and reconstruction, enabling the generation of diverse 3D worlds. Extensive experiments demonstrate that our method achieves state-of-the-art performance in generating coherent, explorable, and interactive 3D worlds while enabling versatile applications in virtual reality, physical simulation, game development, and interactive content creation.

hub tools

citation-role summary

background 4

citation-polarity summary

years

2026 18 2025 1

roles

background 4

polarities

background 4

representative citing papers

XPR: An Extensible Cross-Platform Point-Based Differentiable Renderer

cs.GR · 2026-06-10 · unverdicted · novelty 7.0

XPR is an extensible cross-platform framework for point-based differentiable rendering that decomposes the pipeline into modular operations compilable by XLA, demonstrated with 3DGS, 3DGUT and LinPrim in a few hundred lines of Python each.

EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks

cs.CV · 2026-04-10 · unverdicted · novelty 7.0

EgoTL provides a new egocentric dataset with think-aloud chains and metric labels that benchmarks VLMs on long-horizon tasks and improves their planning, reasoning, and spatial grounding after finetuning.

MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold

cs.CV · 2026-06-11 · unverdicted · novelty 6.0

MoVerse generates real-time interactive video world models from single narrow-FOV images via panoramic diffusion expansion, Gaussian scaffold lifting, and distillation of a bidirectional diffusion teacher into a causal autoregressive renderer.

A Definition and Roadmap for World Models

cs.AI · 2026-07-07 · conditional · novelty 5.0

A perspective article defining world models as finite-resource compression of physical state transitions and outlining a roadmap toward physical AGI via unified representations and interactive simulators.

Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction

cs.CV · 2026-06-11 · unverdicted · novelty 5.0

A multi-view video diffusion model conditioned on relative camera poses via extended RoPE generates dense synchronized views from sparse inputs for 4D Gaussian splatting reconstruction, claiming SOTA results on human datasets and generalization to animals.

InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model

cs.CV · 2026-03-12 · unverdicted · novelty 5.0

InSpatio-WorldFM is a frame-independent generative model that uses explicit 3D anchors and spatial memory to deliver real-time multi-view consistent spatial intelligence via a three-stage training pipeline from pretrained diffusion models.

citing papers explorer

Showing 19 of 19 citing papers.