A multimodal LLM pipeline (Qwen audio + Qwen text embeddings, concatenated and classified) reaches 92.4% accuracy on a combined ADReSS20/ADReSSo21 test set, but the evaluation does not justify state-of-the-art or cross-dataset generalization claims.
hub
Title resolution pending
1 Pith paper cite this work, alongside 3,112 external citations. Polarity classification is still indexing.
1
Pith paper citing it
3,112
external citations · OpenAlex
hub tools
fields
eess.SP 1years
2026 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models
A multimodal LLM pipeline (Qwen audio + Qwen text embeddings, concatenated and classified) reaches 92.4% accuracy on a combined ADReSS20/ADReSSo21 test set, but the evaluation does not justify state-of-the-art or cross-dataset generalization claims.