NUMINA improves counting accuracy in text-to-video diffusion models by up to 7.4% via a training-free identify-then-guide framework on the new CountBench dataset.
Dualdiff+: Dual-branch diffusion for high-fidelity video generation with reward guidance
3 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 3representative citing papers
CityGen is a diffusion-based generative model that synthesizes city-style images from HD maps and visual prompts to enable label-free adaptation for cross-city autonomous driving tasks.
DriVerse is a generative model that simulates driving scenes from an image and trajectory using multimodal prompting and motion alignment, achieving better performance on nuScenes and Waymo datasets with minimal training.
citing papers explorer
-
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
NUMINA improves counting accuracy in text-to-video diffusion models by up to 7.4% via a training-free identify-then-guide framework on the new CountBench dataset.
-
CityGen: Structure-Guided City-Style Synthesis for Cross-City Autonomous Driving
CityGen is a diffusion-based generative model that synthesizes city-style images from HD maps and visual prompts to enable label-free adaptation for cross-city autonomous driving tasks.
-
DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment
DriVerse is a generative model that simulates driving scenes from an image and trajectory using multimodal prompting and motion alignment, achieving better performance on nuScenes and Waymo datasets with minimal training.