Goku provides a 2M-pair dataset for multi-task structural video editing, Goku-Edit model with MLLM and dual-branch design, and Goku-Bench yielding up to 8% gains in instruction following.
In: Proceedings of the AAAI Conference on Artificial Intelligence
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 3years
2026 3representative citing papers
Predicting quantized latent residuals (Latent Drift with FSQ) avoids identity collapse and noise interpolation, improving patient-specific 3D MRI neuro-forecasting over diffusion and autoregressive baselines.
ICDepth adapts text-to-video diffusion transformers for video depth estimation via in-context conditioning, achieving SOTA results on benchmarks with 6-13x less training data than prior generative methods.
citing papers explorer
-
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
Goku provides a 2M-pair dataset for multi-task structural video editing, Goku-Edit model with MLLM and dual-branch design, and Goku-Bench yielding up to 8% gains in instruction following.
-
Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies
Predicting quantized latent residuals (Latent Drift with FSQ) avoids identity collapse and noise interpolation, improving patient-specific 3D MRI neuro-forecasting over diffusion and autoregressive baselines.
-
ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning
ICDepth adapts text-to-video diffusion transformers for video depth estimation via in-context conditioning, achieving SOTA results on benchmarks with 6-13x less training data than prior generative methods.