WMGen-v1 generates diverse long-tail spatial images from one reference image via LVLM scene parsing, LLM-guided expansion, and diffusion synthesis, with detectors trained only on the synthetic data approaching real-data performance on benchmarks.
Lars: A diverse panoptic mar- itime obstacle detection dataset and benchmark,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception
WMGen-v1 generates diverse long-tail spatial images from one reference image via LVLM scene parsing, LLM-guided expansion, and diffusion synthesis, with detectors trained only on the synthetic data approaching real-data performance on benchmarks.