Pith. sign in

REVIEW 2 cited by

When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13276 v2 pith:CHCH72BC submitted 2024-02-17 eess.AS cs.AIcs.SD

classification eess.AScs.AIcs.SD
keywords llmsdepressiondetectionspeechacousticapproachlandmarksmental
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Depression is a critical concern in global mental health, prompting extensive research into AI-based detection methods. Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in mental healthcare applications. However, their primary limitation arises from their exclusive dependence on textual input, which constrains their overall capabilities. Furthermore, the utilization of LLMs in identifying and analyzing depressive states is still relatively untapped. In this paper, we present an innovative approach to integrating acoustic speech information into the LLMs framework for multimodal depression detection. We investigate an efficient method for depression detection by integrating speech signals into LLMs utilizing Acoustic Landmarks. By incorporating acoustic landmarks, which are specific to the pronunciation of spoken words, our method adds critical dimensions to text transcripts. This integration also provides insights into the unique speech patterns of individuals, revealing the potential mental states of individuals. Evaluations of the proposed approach on the DAIC-WOZ dataset reveal state-of-the-art results when compared with existing Audio-Text baselines. In addition, this approach is not only valuable for the detection of depression but also represents a new perspective in enhancing the ability of LLMs to comprehend and process speech signals.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CASPER: A Large Scale Spontaneous Speech Dataset

    cs.CL 2025-05 conditional novelty 6.0 of 10

    CASPER presents a 102-hour spontaneous English conversation dataset with per-speaker channels, speaker metadata, and baseline ASR and diarization results.

  2. Speech as a Multimodal Digital Phenotype for Multi-Task LLM-based Mental Health Prediction

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A trimodal, longitudinal, multi-task LLM pipeline predicts adolescent depression with 70.8% balanced accuracy on the private DEW dataset, but the gain over simpler baselines is modest and lacks external validation.

Pith tools