Pith. sign in

REVIEW 3 cited by

DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09345 v1 pith:QXFA426C submitted 2024-06-13 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords speechencoderlanguageself-supervisedspokentasksadapteranswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The integration of pre-trained text-based large language models (LLM) with speech input has enabled instruction-following capabilities for diverse speech tasks. This integration requires the use of a speech encoder, a speech adapter, and an LLM, trained on diverse tasks. We propose the use of discrete speech units (DSU), rather than continuous-valued speech encoder outputs, that are converted to the LLM token embedding space using the speech adapter. We generate DSU using a self-supervised speech encoder followed by k-means clustering. The proposed model shows robust performance on speech inputs from seen/unseen domains and instruction-following capability in spoken question answering. We also explore various types of DSU extracted from different layers of the self-supervised speech encoder, as well as Mel frequency Cepstral Coefficients (MFCC). Our findings suggest that the ASR task and datasets are not crucial in instruction-tuning for spoken question answering tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Quantizing speech at frame, phone, word, and utterance levels preserves emotion and prominence better than frame-only discrete units at similar bitrates.

  2. The Role of Prosody in Spoken Question Answering

    cs.CL 2025-02 conditional novelty 6.0 of 10

    On natural-speech spoken QA, prosodic-only models beat chance but trail lexical models, and lexical cues dominate whenever both are available.

  3. Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    Under matched settings, continuous SSL speech features generally outperform discrete tokens on six spoken language understanding tasks in SpeechLLMs.

Pith tools