Pith. sign in

REVIEW 35 cited by

FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04051 v3 pith:M5PNVHOT submitted 2024-07-04 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords voicemodelsfunaudiollmcosyvoicegenerationgithublanguagesllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative models: SenseVoice, which handles multilingual speech recognition, emotion recognition, and audio event detection; and CosyVoice, which facilitates natural speech generation with control over multiple languages, timbre, speaking style, and speaker identity. SenseVoice-Small delivers exceptionally low-latency ASR for 5 languages, and SenseVoice-Large supports high-precision ASR for over 50 languages, while CosyVoice excels in multi-lingual voice generation, zero-shot in-context learning, cross-lingual voice cloning, and instruction-following capabilities. The models related to SenseVoice and CosyVoice have been open-sourced on Modelscope and Huggingface, along with the corresponding training, inference, and fine-tuning codes released on GitHub. By integrating these models with LLMs, FunAudioLLM enables applications such as speech-to-speech translation, emotional voice chat, interactive podcasts, and expressive audiobook narration, thereby pushing the boundaries of voice interaction technology. Demos are available at https://fun-audio-llm.github.io, and the code can be accessed at https://github.com/FunAudioLLM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 35 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise

    cs.IR 2026-02 unverdicted novelty 7.0 of 10

    SQuTR is a large bilingual benchmark of 37,317 synthesized spoken queries under clean/low/medium/high noise, showing that retrieval quality steadily degrades as noise increases.

  2. Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A new spoken math benchmark, Spoken-MQA, shows that current speech-based AI models reason poorly from spoken math input, especially for arithmetic and knowledge-heavy problems.

  3. A Non-autoregressive Model for Joint STT and TTS

    cs.SD 2025-01 conditional novelty 7.0 of 10

    A joint non-autoregressive model handles both STT and TTS in one framework, beating its own STT baseline and matching its TTS baseline with extra unpaired data and iterative refinement.

  4. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  5. Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis

    cs.HC 2026-07 conditional novelty 6.0 of 10

    AU-supervised single-token face encoding plus dual visual–speech DPO on a large real-conversation dataset improves empathetic conversational TTS over text/speech-only and prior visual CSS systems.

  6. Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Lychee-FD resolves modality interference in full-duplex spoken language models by separating acoustic and semantic parameters in deep layers and adding a dense semantic alignment channel, achieving state-of-the-art pe...

  7. OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    OmniHuman is a new large-scale multi-scene dataset with video-, frame-, and individual-level annotations for human-centric video generation, accompanied by the OHBench benchmark that adds metrics aligned with human pe...

  8. A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff

    cs.IT 2026-04 unverdicted novelty 6.0 of 10

    Synset-based reconstruction and synonymous variational inference are claimed to derive the distributional divergence in RDP and unify it with classical rate-distortion theory.

  9. Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A unified mask-based discrete diffusion model jointly models text, speech, and image tokens and matches or exceeds several any-to-any and specialist multimodal baselines.

  10. Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards

    cs.SD 2026-01 conditional novelty 6.0 of 10

    Semantic-token infilling plus a frozen flow-matching decoder and a TTS-based GRPO reward produces more intelligible and natural text-based speech edits than prior AR and NAR baselines.

  11. The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Text jailbreak prompts converted to audio match or beat dedicated audio jailbreaks on omni-models, and transfer success tracks how tightly the model aligns text and audio representations.

  12. AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues

    cs.MM 2025-10 conditional novelty 6.0 of 10

    Current omni-modal LLMs underperform on audio-visual emotional reasoning, and automatic scores diverge from human perceptual judgments; AV-EMO-Reasoning provides a benchmark to measure this.

  13. WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    The authors built and released the largest open-source Cantonese speech corpus (21,800 hours, 10 domains, rich metadata), and show that models trained on it match or beat existing speech recognition and synthesis systems.

  14. NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

    cs.SD 2025-08 unverdicted novelty 6.0 of 10

    A Mandarin speech pipeline that tags non-word vocalizations (laughter, breath, filled pauses) for ASR and controls their generation in TTS, backed by a claimed 573-hour, 174,179-utterance word-level annotated corpus.

  15. Differentiable Reward Optimization for LLM based TTS system

    cs.SD 2025-07 conditional novelty 6.0 of 10

    DiffRO optimizes codec-based TTS models directly on differentiable token-level rewards, improving WER and enabling zero-shot emotion control.

  16. DEBATE: A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech

    cs.CL 2025-06 conditional novelty 6.0 of 10

    DEBATE is a first-of-its-kind Mandarin speech-text dataset for studying how prosody resolves textual ambiguity, and current large speech-language models lag well behind humans on it.

  17. Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Compressed-to-fine language modeling improves speech token prediction by retaining prompt and local tokens while compressing long-range token spans into compact summaries.

  18. Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Chain-of-Talkers (CoTalk) has annotators sequentially dictate only the missing visual details, and it reports modest gains in annotation speed and caption density over parallel typed annotation.

  19. Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models

    cs.SD 2025-05 conditional novelty 6.0 of 10

    AJailBench is an open benchmark showing that large audio-language models can be jailbroken through TTS-converted text attacks and through subtle acoustic perturbations that preserve speech semantics.

  20. CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech

    cs.SD 2025-01 conditional novelty 6.0 of 10

    A new large-scale dataset and codec taxonomy show that codec re-synthesized speech, especially balanced by decoder type, trains detectors that catch codec-based deepfake speech better than traditional anti-spoofing training.

  21. DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions

    cs.SD 2025-01 conditional novelty 6.0 of 10

    A sketch-conditioned diffusion TTS model that turns coarse user-drawn pitch and energy trends into natural, precisely controlled speech.

  22. OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

    cs.CL 2025-01 conditional novelty 6.0 of 10

    ShareChatX and OmniChat show that large-scale synthetic spoken dialogue data improves multi-turn response quality and emotion prediction, setting a new state of the art on DailyTalk.

  23. X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

    cs.LG 2026-07 conditional novelty 5.0 of 10

    X3-OPD improves audio-grounded reasoning by training the audio student on its own rollouts with token-level teacher feedback, using a three-tier paired text-audio corpus.

  24. X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

    eess.AS 2026-07 conditional novelty 5.0 of 10

    An open, modular cascaded system (streaming ASR + MT + prompt-conditioned TTS) preserves speaker identity in long-form multi-speaker translation, at higher latency and slightly lower translation quality than proprietary APIs.

  25. CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays

    cs.SD 2025-09 conditional novelty 5.0 of 10

    CabinSep cuts in-car ASR character error by 17.5% relative to DualSep with a 0.4 GMACs mask-based MVDR system trained on mixed simulated and real impulse responses.

  26. ChipChat: Low-Latency Cascaded Conversational Agent in MLX

    eess.AS 2025-08 conditional novelty 5.0 of 10

    ChipChat is a fully on-device cascaded voice agent that reports about 920 ms total latency using streaming ASR, a state-action LLM, streaming TTS, and a vocoder.

  27. Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...

  28. SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset

    cs.CL 2025-05 reject novelty 5.0 of 10

    The authors present SwitchLingua, a large multilingual code-switching text and audio dataset, and SAER, a semantic-aware error metric for code-switching ASR evaluation.

  29. Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection

    cs.SD 2025-05 conditional novelty 5.0 of 10

    An LLM-driven multi-task system reports a 5.45% CER and 73.63% average SED F1 on the AS-70 Mandarin stuttering benchmark, though key baselines and uncertainty are missing.

  30. A Large-scale Dataset with Behavior, Attributes, and Content of Mobile Short-video Platform

    cs.MM 2025-02 conditional novelty 5.0 of 10

    A new dataset combines short-video user behavior, user attributes, and video content at a scale larger than most public benchmarks.

  31. OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

    cs.SD 2025-01 conditional novelty 5.0 of 10

    An open, resource-lean speech understanding LLM trained on 50,500 hours matches or beats larger industry models on several Chinese benchmarks, with caveats in its internal evaluation.

  32. MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A single aligned multimodal model, MinMo, achieves strong or state-of-the-art results on speech recognition, translation, emotion recognition, voice generation, and full-duplex dialogue while preserving the underlying...

  33. VibeVoice-ASR-BitNet Technical Report

    cs.SD 2026-07 conditional novelty 4.0 of 10

    Heterogeneous quantization (INT8 tokenizer + 2-bit ternary LM) makes a 1.5B-parameter LLM-based ASR system run at real-time speed on CPUs with 2.9x compression and modest measured WER increases.

  34. Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge

    cs.SD 2025-05 conditional novelty 4.0 of 10

    A hybrid diarization and ASR system with a CER-supervised bridging module achieved the best results in two MISP 2025 tracks.

  35. HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

    cs.CV 2025-01 reject novelty 4.0 of 10

    A human-centric vision-speech language model with three instruction-weighted visual branches and audio input reports strong emotion, facial expression, and action results, but its evaluation is under-specified and not...

Pith tools