CookVoice is a single compact model that generates both speech and singing, aligning content, style, and prosody at the frame level and reporting stronger style and pitch controllability than larger unified baselines.
This section will present the architecture of CookV oice in detail
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
CookVoice is a single compact model that generates both speech and singing, aligning content, style, and prosody at the frame level and reporting stronger style and pitch controllability than larger unified baselines.