Pith. sign in

REVIEW 3 cited by

PsyEval: A Suite of Mental Health Related Tasks for Evaluating Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09189 v2 pith:XOY2A5CH submitted 2023-11-15 cs.CL

classification cs.CL
keywords mentalpsyevalevaluatinghealthllmstaskscomprehensivedomain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Evaluating Large Language Models (LLMs) in the mental health domain poses distinct challenged from other domains, given the subtle and highly subjective nature of symptoms that exhibit significant variability among individuals. This paper presents PsyEval, the first comprehensive suite of mental health-related tasks for evaluating LLMs. PsyEval encompasses five sub-tasks that evaluate three critical dimensions of mental health. This comprehensive framework is designed to thoroughly assess the unique challenges and intricacies of mental health-related tasks, making PsyEval a highly specialized and valuable tool for evaluating LLM performance in this domain. We evaluate twelve advanced LLMs using PsyEval. Experiment results not only demonstrate significant room for improvement in current LLMs concerning mental health but also unveil potential directions for future model optimization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The methodology of Constructing the Large-Scale Dataset for Detecting Presuicidal and Anti-Suicidal Signals in Social Media Texts in Russian

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A new methodology and a 57,810-example Russian social media dataset with 33 presuicidal and 12 anti-suicidal classes, plus baseline RuBERT experiments.

  2. LLMs vs. Chinese Anime Enthusiasts: A Comparative Study on Emotionally Supportive Role-Playing

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    ChatAnime, a new emotionally supportive anime role-play benchmark, reports top LLMs outperforming human enthusiasts on role-playing and emotional support metrics while humans keep the diversity edge.

  3. Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    OpenAI o1 and DeepSeek-R1, evaluated on a new synthetic multi-turn mental-health dialogue benchmark, scored below 'good' on patient-centric communication and only about 31% on exact diagnosis matching.

Pith tools