Pith. sign in

REVIEW 27 cited by

LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.04673 v4 pith:NRK3Z5UV submitted 2023-10-07 cs.SD cs.AIcs.LGcs.MMeess.AS

classification cs.SDcs.AIcs.LGcs.MMeess.AS
keywords audiolauragptspeechrecognitionaudio-and-textcodeccontinuousdiscrete
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for audio-and-text large language models (LLMs). Previous mainstream audio-and-text LLMs use discrete audio tokens to represent both input and output audio; however, they suffer from performance degradation on tasks such as automatic speech recognition, speech-to-text translation, and speech enhancement over models using continuous speech features. In this paper, we propose LauraGPT, a novel unified audio-and-text GPT-based LLM for audio recognition, understanding, and generation. LauraGPT is a versatile LLM that can process both audio and text inputs and generate outputs in either modalities. We propose a novel data representation that combines continuous and discrete features for audio: LauraGPT encodes input audio into continuous representations using an audio encoder and generates output audio from discrete codec codes. We propose a one-step codec vocoder to overcome the prediction challenge caused by the multimodal distribution of codec tokens. We fine-tune LauraGPT using supervised multi-task learning. Extensive experiments show that LauraGPT consistently achieves comparable to superior performance compared to strong baselines on a wide range of audio tasks related to content, semantics, paralinguistics, and audio-signal analysis, such as automatic speech recognition, speech-to-text translation, text-to-speech synthesis, speech enhancement, automated audio captioning, speech emotion recognition, and spoken language understanding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Non-autoregressive Model for Joint STT and TTS

    cs.SD 2025-01 conditional novelty 7.0 of 10

    A joint non-autoregressive model handles both STT and TTS in one framework, beating its own STT baseline and matching its TTS baseline with extra unpaired data and iterative refinement.

  2. FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech

    eess.AS 2026-06 unverdicted novelty 6.0 of 10

    FlowTTS-GRPO fine-tunes open-source flow-matching TTS models with multi-objective online RL via ODE-to-SDE conversion, improving speaker similarity and quality on CosyVoice 3.0 and F5-TTS.

  3. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  4. CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

    cs.SD 2025-09 unverdicted novelty 6.0 of 10

    CodecSep performs prompt-driven universal sound separation directly in neural audio codec latents by combining a frozen DAC backbone with a lightweight FiLM-conditioned Transformer masker driven by CLAP embeddings, yi...

  5. LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.

  6. Self-Improvement for Audio Large Language Model using Unlabeled Speech

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SI-SDA uses attention-saliency prompt-reliance as an unsupervised reward to reinforce a model's own beam hypotheses, improving audio LLM WER and BLEU across ASR, S2TT, and SQA without labeled data.

  7. A Variational Framework for Improving Naturalness in Generative Spoken Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Integrating a VAE with an autoregressive prior into token-based spoken language modeling learns continuous variational features that improve naturalness without hand-engineered pitch, with human raters preferring the ...

  8. VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.

  9. Towards Reliable Large Audio Language Model

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Training a large audio language model to say 'I don't know' on one audio type (speech, music, or sound) makes it more likely to refuse uncertain questions on the other types.

  10. CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech

    cs.SD 2025-01 conditional novelty 6.0 of 10

    A new large-scale dataset and codec taxonomy show that codec re-synthesized speech, especially balanced by decoder type, trains detectors that catch codec-based deepfake speech better than traditional anti-spoofing training.

  11. TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

    cs.SD 2024-12 conditional novelty 6.0 of 10

    TouchTTS reports a 51.6% data retention rate using a noise-robust tokenizer and two-ASR cross-validation, and a Qwen-backbone flow model that unifies streaming and non-streaming synthesis while matching CosyVoice on PER.

  12. A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    Across six speech-understanding tasks, continuous SSL features beat k-means discrete tokens in almost all cases when paired with a 0.5B instruction-tuned LLM.

  13. FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

    cs.SD 2025-09 conditional novelty 5.0 of 10

    A full-duplex voice system with streaming personalized VAD and semantic end-of-turn detection reports fewer false barge-ins and latencies near commercial benchmarks.

  14. Group Relative Policy Optimization for Speech Recognition

    eess.AS 2025-09 conditional novelty 5.0 of 10

    Applying GRPO with rule-based rewards to LLM-based ASR improves WER by up to 18.4% relative and reduces hallucination errors on unseen acoustic conditions.

  15. Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection

    cs.SD 2025-05 conditional novelty 5.0 of 10

    An LLM-driven multi-task system reports a 5.45% CER and 73.63% average SED F1 on the AS-70 Mandarin stuttering benchmark, though key baselines and uncertainty are missing.

  16. Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A codec-token denoiser that predicts only the first two token groups of clean audio enables noise-robust zero-shot TTS, matching clean-prompt performance.

  17. Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A pipeline condenses in-the-wild speech with emotion labels and uses ChatGPT to generate contextual paralinguistic QA pairs, released as a 480-sample benchmark.

  18. Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    Under matched settings, continuous SSL speech features generally outperform discrete tokens on six spoken language understanding tasks in SpeechLLMs.

  19. BoSS: Beyond-Semantic Speech

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Current spoken-language models perform poorly on a new five-task evaluation of beyond-semantic speech signals, including dialect, emotion, age, and non-verbal cues.

  20. Breaking the Barriers of Text-Hungry and Audio-Deficient AI

    cs.SD 2025-06 reject novelty 4.0 of 10

    A proposed audio-native translation framework called MAST with fractional diffusion is described, but no evidence is given that it produces working translations.

  21. Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget

    cs.SD 2025-04 conditional novelty 4.0 of 10

    Muyan-TTS, a 3B-parameter LLM-based TTS model trained on 100,000+ hours of podcast audio, produces competitive zero-shot speech and runs at 0.33 seconds of inference per second of speech.

  22. SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization

    cs.CV 2024-12 reject novelty 4.0 of 10

    SilVar fuses speech, image, and text via CLIP, Whisper, and LLaMA to perform reasoning visual question answering and object localization, but its claimed state-of-the-art results are not supported by its own tables.

  23. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.

  24. YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

    cs.SD 2024-12 reject novelty 4.0 of 10

    A video-guided sound effects model with a learnable audio-visual aggregator and multi-modal chain-of-thought module reports strong VGGSound benchmark scores, but its few-shot claim rests on three qualitative samples.

  25. Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding

    eess.AS 2024-12 reject novelty 4.0 of 10

    A modular audio chatbot using a BERT intent router, expert audio models, and a 3.8B LLM over audio-event metadata matches 7B-8B audio-language models on MMAU sound and beats several of them on custom temporal QA.

  26. AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM

    eess.AS 2024-12 conditional novelty 4.0 of 10

    AlignFormer, a CTC-guided dynamic-window adapter, lets a frozen LLM trained on ASR data alone achieve high instruction-following rates on zero-shot speech translation and question answering.

  27. WavChat: A Survey of Spoken Dialogue Models

    eess.AS 2024-11 conditional novelty 4.0 of 10

    WavChat categorizes spoken dialogue models into cascaded and end-to-end paradigms and surveys speech representations, training strategies, streaming, duplex interaction, datasets, and evaluation benchmarks.

Pith tools