REVIEW 3 cited by
PsyEval: A Suite of Mental Health Related Tasks for Evaluating Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Evaluating Large Language Models (LLMs) in the mental health domain poses distinct challenged from other domains, given the subtle and highly subjective nature of symptoms that exhibit significant variability among individuals. This paper presents PsyEval, the first comprehensive suite of mental health-related tasks for evaluating LLMs. PsyEval encompasses five sub-tasks that evaluate three critical dimensions of mental health. This comprehensive framework is designed to thoroughly assess the unique challenges and intricacies of mental health-related tasks, making PsyEval a highly specialized and valuable tool for evaluating LLM performance in this domain. We evaluate twelve advanced LLMs using PsyEval. Experiment results not only demonstrate significant room for improvement in current LLMs concerning mental health but also unveil potential directions for future model optimization.
Forward citations
Cited by 3 Pith papers
-
The methodology of Constructing the Large-Scale Dataset for Detecting Presuicidal and Anti-Suicidal Signals in Social Media Texts in Russian
A new methodology and a 57,810-example Russian social media dataset with 33 presuicidal and 12 anti-suicidal classes, plus baseline RuBERT experiments.
-
LLMs vs. Chinese Anime Enthusiasts: A Comparative Study on Emotionally Supportive Role-Playing
ChatAnime, a new emotionally supportive anime role-play benchmark, reports top LLMs outperforming human enthusiasts on role-playing and emotional support metrics while humans keep the diversity edge.
-
Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations
OpenAI o1 and DeepSeek-R1, evaluated on a new synthetic multi-turn mental-health dialogue benchmark, scored below 'good' on patient-centric communication and only about 31% on exact diagnosis matching.
Discussion (0). Continue with ORCID to comment.