MiDashengLM-Gen uses an LLM backbone with per-token flow matching to generate variable-length multilingual audio scenes with near-TTS speech intelligibility and competitive mixed-scene quality.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching
MiDashengLM-Gen uses an LLM backbone with per-token flow matching to generate variable-length multilingual audio scenes with near-TTS speech intelligibility and competitive mixed-scene quality.