JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

· 2026 · cs.GR · arXiv 2605.04128

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open full Pith review browse 1 citing papers arXiv PDF

abstract

We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing. JoyAI-Image couples a spatially enhanced Multimodal Large Language Model (MLLM) with a Multimodal Diffusion Transformer (MMDiT), allowing perception and generation to interact through a shared multimodal interface. Around this architecture, we build a scalable training recipe that combines unified instruction tuning, long-text rendering supervision, spatially grounded data, and both general and spatial editing signals. This design gives the model broad multimodal capability while strengthening geometry-aware reasoning and controllable visual synthesis. Experiments across understanding, generation, long-text rendering, and editing benchmarks show that JoyAI-Image achieves state-of-the-art or highly competitive performance. More importantly, the bidirectional loop between enhanced understanding, controllable spatial editing, and novel-view-assisted reasoning enables the model to move beyond general visual competence toward stronger spatial intelligence. These results suggest a promising path for unified visual models in downstream applications such as vision-language-action systems and world models.

representative citing papers

Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models

cs.CV · 2026-05-27 · unverdicted · novelty 7.0

Embodied3DBench creates a new evaluation benchmark for low-level embodied spatial intelligence in VLMs, evaluates 13 models showing gaps in interaction perception, and supplies a large synthetic training set that yields measurable gains.

citing papers explorer

Showing 1 of 1 citing paper after filters.

Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models cs.CV · 2026-05-27 · unverdicted · none · ref 46 · internal anchor
Embodied3DBench creates a new evaluation benchmark for low-level embodied spatial intelligence in VLMs, evaluates 13 models showing gaps in interaction perception, and supplies a large synthetic training set that yields measurable gains.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

fields

years

verdicts

representative citing papers

citing papers explorer