Mini-Gemini enhances VLMs via high-resolution visual refinement, curated reasoning data, and self-guided generation to reach leading zero-shot benchmark results across 2B-34B LLMs.
Llmga: Multimodal large language model based generation assistant
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 2verdicts
UNVERDICTED 2representative citing papers
WMGen-v1 generates diverse long-tail spatial images from one reference image via LVLM scene parsing, LLM-guided expansion, and diffusion synthesis, with detectors trained only on the synthetic data approaching real-data performance on benchmarks.
citing papers explorer
-
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Mini-Gemini enhances VLMs via high-resolution visual refinement, curated reasoning data, and self-guided generation to reach leading zero-shot benchmark results across 2B-34B LLMs.
-
One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception
WMGen-v1 generates diverse long-tail spatial images from one reference image via LVLM scene parsing, LLM-guided expansion, and diffusion synthesis, with detectors trained only on the synthetic data approaching real-data performance on benchmarks.