Pith. sign in

REVIEW 5 cited by

MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.13509 v2 pith:MYB7JFOD submitted 2025-03-13 cs.LG cs.AIcs.CLcs.CYcs.HC

classification cs.LGcs.AIcs.CLcs.CYcs.HC
keywords datasetmentalchat16khealthmentalassistancebenchmarkconversationalgithub
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce MentalChat16K, an English benchmark dataset combining a synthetic mental health counseling dataset and a dataset of anonymized transcripts from interventions between Behavioral Health Coaches and Caregivers of patients in palliative or hospice care. Covering a diverse range of conditions like depression, anxiety, and grief, this curated dataset is designed to facilitate the development and evaluation of large language models for conversational mental health assistance. By providing a high-quality resource tailored to this critical domain, MentalChat16K aims to advance research on empathetic, personalized AI solutions to improve access to mental health support services. The dataset prioritizes patient privacy, ethical considerations, and responsible data usage. MentalChat16K presents a valuable opportunity for the research community to innovate AI technologies that can positively impact mental well-being. The dataset is available at https://huggingface.co/datasets/ShenLab/MentalChat16K and the code and documentation are hosted on GitHub at https://github.com/ChiaPatricia/MentalChat16K.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis

    cs.MA 2026-02 conditional novelty 7.0 of 10

    A 16,000-case Chinese psychiatric consultation benchmark shows LLMs reach ~92% accuracy on depression-vs-anxiety but only ~29–43% on comorbidity and 12-way differential diagnosis, and dynamic interviewing does not rel...

  2. A Comprehensive Review of Datasets for Clinical Mental Health AI Systems

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A systematic catalog of 89 clinical mental health datasets and 16 synthetic datasets, with a gap analysis on access, culture, and modality.

  3. EmoStage: A Framework for Accurate Empathetic Response Generation via Perspective-Taking and Phase Recognition

    cs.CL 2025-06 conditional novelty 5.0 of 10

    EmoStage improves LLM counseling responses by prompting models to first take the client's perspective and recognize the counseling stage, with no training data.

  4. RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being

    cs.AI 2025-06 reject novelty 5.0 of 10

    Structurally prompted health LLM responses score higher than zero-shot and few-shot prompting on AI-judged safety and quality metrics across four well-being datasets.

  5. Trusting What You Cannot See: Auditable Fine-Tuning and Inference for Proprietary AI

    cs.CR 2026-03 reject novelty 3.0 of 10

    AFTUNE spot-checks cloud fine-tuning and inference by hashing boundary states and recomputing sampled blocks inside a TEE, but its detection-probability formula assumes tampered blocks are detectable when sampled; int...

Pith tools