REVIEW 35 cited by
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative models: SenseVoice, which handles multilingual speech recognition, emotion recognition, and audio event detection; and CosyVoice, which facilitates natural speech generation with control over multiple languages, timbre, speaking style, and speaker identity. SenseVoice-Small delivers exceptionally low-latency ASR for 5 languages, and SenseVoice-Large supports high-precision ASR for over 50 languages, while CosyVoice excels in multi-lingual voice generation, zero-shot in-context learning, cross-lingual voice cloning, and instruction-following capabilities. The models related to SenseVoice and CosyVoice have been open-sourced on Modelscope and Huggingface, along with the corresponding training, inference, and fine-tuning codes released on GitHub. By integrating these models with LLMs, FunAudioLLM enables applications such as speech-to-speech translation, emotional voice chat, interactive podcasts, and expressive audiobook narration, thereby pushing the boundaries of voice interaction technology. Demos are available at https://fun-audio-llm.github.io, and the code can be accessed at https://github.com/FunAudioLLM.
Forward citations
Cited by 35 Pith papers
-
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise
SQuTR is a large bilingual benchmark of 37,317 synthesized spoken queries under clean/low/medium/high noise, showing that retrieval quality steadily degrades as noise increases.
-
Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
A new spoken math benchmark, Spoken-MQA, shows that current speech-based AI models reason poorly from spoken math input, especially for arithmetic and knowledge-heavy problems.
-
A Non-autoregressive Model for Joint STT and TTS
A joint non-autoregressive model handles both STT and TTS in one framework, beating its own STT baseline and matching its TTS baseline with extra unpaired data and iterative refinement.
-
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.
-
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
AU-supervised single-token face encoding plus dual visual–speech DPO on a large real-conversation dataset improves empathetic conversational TTS over text/speech-only and prior visual CSS systems.
-
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
Lychee-FD resolves modality interference in full-duplex spoken language models by separating acoustic and semantic parameters in deep layers and adding a dense semantic alignment channel, achieving state-of-the-art pe...
-
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
OmniHuman is a new large-scale multi-scene dataset with video-, frame-, and individual-level annotations for human-centric video generation, accompanied by the OHBench benchmark that adds metrics aligned with human pe...
-
A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff
Synset-based reconstruction and synonymous variational inference are claimed to derive the distributional divergence in RDP and unify it with classical rate-distortion theory.
-
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
A unified mask-based discrete diffusion model jointly models text, speech, and image tokens and matches or exceeds several any-to-any and specialist multimodal baselines.
-
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
Semantic-token infilling plus a frozen flow-matching decoder and a TTS-based GRPO reward produces more intelligible and natural text-based speech edits than prior AR and NAR baselines.
-
The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer
Text jailbreak prompts converted to audio match or beat dedicated audio jailbreaks on omni-models, and transfer success tracks how tightly the model aligns text and audio representations.
-
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
Current omni-modal LLMs underperform on audio-visual emotional reasoning, and automatic scores diverge from human perceptual judgments; AV-EMO-Reasoning provides a benchmark to measure this.
-
WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation
The authors built and released the largest open-source Cantonese speech corpus (21,800 hours, 10 domains, rich metadata), and show that models trained on it match or beat existing speech recognition and synthesis systems.
-
NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations
A Mandarin speech pipeline that tags non-word vocalizations (laughter, breath, filled pauses) for ASR and controls their generation in TTS, backed by a claimed 573-hour, 174,179-utterance word-level annotated corpus.
-
Differentiable Reward Optimization for LLM based TTS system
DiffRO optimizes codec-based TTS models directly on differentiable token-level rewards, improving WER and enabling zero-shot emotion control.
-
DEBATE: A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech
DEBATE is a first-of-its-kind Mandarin speech-text dataset for studying how prosody resolves textual ambiguity, and current large speech-language models lag well behind humans on it.
-
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
Compressed-to-fine language modeling improves speech token prediction by retaining prompt and local tokens while compressing long-range token spans into compact summaries.
-
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
Chain-of-Talkers (CoTalk) has annotators sequentially dictate only the missing visual details, and it reports modest gains in annotation speed and caption density over parallel typed annotation.
-
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
AJailBench is an open benchmark showing that large audio-language models can be jailbroken through TTS-converted text attacks and through subtle acoustic perturbations that preserve speech semantics.
-
CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech
A new large-scale dataset and codec taxonomy show that codec re-synthesized speech, especially balanced by decoder type, trains detectors that catch codec-based deepfake speech better than traditional anti-spoofing training.
-
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
A sketch-conditioned diffusion TTS model that turns coarse user-drawn pitch and energy trends into natural, precisely controlled speech.
-
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
ShareChatX and OmniChat show that large-scale synthetic spoken dialogue data improves multi-turn response quality and emotion prediction, setting a new state of the art on DailyTalk.
-
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
X3-OPD improves audio-grounded reasoning by training the audio student on its own rollouts with token-level teacher feedback, using a three-tier paired text-audio corpus.
-
X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System
An open, modular cascaded system (streaming ASR + MT + prompt-conditioned TTS) preserves speaker identity in long-form multi-speaker translation, at higher latency and slightly lower translation quality than proprietary APIs.
-
CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays
CabinSep cuts in-car ASR character error by 17.5% relative to DualSep with a 0.4 GMACs mask-based MVDR system trained on mixed simulated and real impulse responses.
-
ChipChat: Low-Latency Cascaded Conversational Agent in MLX
ChipChat is a fully on-device cascaded voice agent that reports about 920 ms total latency using streaming ASR, a state-action LLM, streaming TTS, and a vocoder.
-
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...
-
SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
The authors present SwitchLingua, a large multilingual code-switching text and audio dataset, and SAER, a semantic-aware error metric for code-switching ASR evaluation.
-
Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection
An LLM-driven multi-task system reports a 5.45% CER and 73.63% average SED F1 on the AS-70 Mandarin stuttering benchmark, though key baselines and uncertainty are missing.
-
A Large-scale Dataset with Behavior, Attributes, and Content of Mobile Short-video Platform
A new dataset combines short-video user behavior, user attributes, and video content at a scale larger than most public benchmarks.
-
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
An open, resource-lean speech understanding LLM trained on 50,500 hours matches or beats larger industry models on several Chinese benchmarks, with caveats in its internal evaluation.
-
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
A single aligned multimodal model, MinMo, achieves strong or state-of-the-art results on speech recognition, translation, emotion recognition, voice generation, and full-duplex dialogue while preserving the underlying...
-
VibeVoice-ASR-BitNet Technical Report
Heterogeneous quantization (INT8 tokenizer + 2-bit ternary LM) makes a 1.5B-parameter LLM-based ASR system run at real-time speed on CPUs with 2.9x compression and modest measured WER increases.
-
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
A hybrid diarization and ASR system with a CER-supervised bridging module achieved the best results in two MISP 2025 tracks.
-
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
A human-centric vision-speech language model with three instruction-weighted visual branches and audio input reports strong emotion, facial expression, and action results, but its evaluation is under-specified and not...
Discussion (0). Continue with ORCID to comment.