A multimodal LLM pipeline with Wav2Vec audio and WHO-based Q&A knowledge injection reports small improvements on DAIC-WOZ, but the fusion adds nothing over audio-only and the baseline is cherry-picked.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.HC 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Large Language Models for Depression Recognition in Spoken Language Integrating Psychological Knowledge
A multimodal LLM pipeline with Wav2Vec audio and WHO-based Q&A knowledge injection reports small improvements on DAIC-WOZ, but the fusion adds nothing over audio-only and the baseline is cherry-picked.