A single canonical neural atlas is learned jointly over thousands of ultrasound frames from five cardiac and musculoskeletal datasets via DINOv3 features and per-video generative latent optimization embeddings to support annotation transfer.
SVG- T2I: Scaling up text-to-image latent diffusion model without variational autoencoder
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 4years
2026 4roles
background 2polarities
background 2representative citing papers
A 5B-parameter text-to-image model trained in 125K A800 GPU hours matches or beats much more expensive models on several benchmarks via semantic-first asynchronous diffusion.
EponaV2 advances perception-free driving world models by forecasting comprehensive future 3D geometry and semantic representations, achieving SOTA planning performance on NAVSIM benchmarks.
Using understanding tasks as direct supervision during post-training improves image generation and editing in unified multimodal models.
citing papers explorer
-
Cohort-Scale Neural Atlases of Ultrasound Video
A single canonical neural atlas is learned jointly over thousands of ultrasound frames from five cardiac and musculoskeletal datasets via DINOv3 features and per-video generative latent optimization embeddings to support annotation transfer.
-
SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion
A 5B-parameter text-to-image model trained in 125K A800 GPU hours matches or beats much more expensive models on several benchmarks via semantic-first asynchronous diffusion.
-
EponaV2: Driving World Model with Comprehensive Future Reasoning
EponaV2 advances perception-free driving world models by forecasting comprehensive future 3D geometry and semantic representations, achieving SOTA planning performance on NAVSIM benchmarks.
-
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
Using understanding tasks as direct supervision during post-training improves image generation and editing in unified multimodal models.