MAGUS orchestrates a single multimodal LLM and several diffusion models in a two-phase agent pipeline with confidence-thresholded beam search, reporting modest gains over its own base model on vision, video, audio, and generation benchmarks.
InForty-first international conference on machine learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation
MAGUS orchestrates a single multimodal LLM and several diffusion models in a two-phase agent pipeline with confidence-thresholded beam search, reporting modest gains over its own base model on vision, video, audio, and generation benchmarks.