Pith. sign in

REVIEW 2 cited by

Underneath the Numbers: Quantitative and Qualitative Gender Fairness in LLMs for Depression Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08183 v2 pith:VYKXJJCS submitted 2024-06-12 cs.CL

classification cs.CL
keywords llmsevaluationfairnessqualitativebiasquantitativechatgptdepression
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent studies show bias in many machine learning models for depression detection, but bias in LLMs for this task remains unexplored. This work presents the first attempt to investigate the degree of gender bias present in existing LLMs (ChatGPT, LLaMA 2, and Bard) using both quantitative and qualitative approaches. From our quantitative evaluation, we found that ChatGPT performs the best across various performance metrics and LLaMA 2 outperforms other LLMs in terms of group fairness metrics. As qualitative fairness evaluation remains an open research question we propose several strategies (e.g., word count, thematic analysis) to investigate whether and how a qualitative evaluation can provide valuable insights for bias analysis beyond what is possible with quantitative evaluation. We found that ChatGPT consistently provides a more comprehensive, well-reasoned explanation for its prediction compared to LLaMA 2. We have also identified several themes adopted by LLMs to qualitatively evaluate gender fairness. We hope our results can be used as a stepping stone towards future attempts at improving qualitative evaluation of fairness for LLMs especially for high-stakes tasks such as depression detection.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    Zero-shot vision-language models are unreliable and vary widely for depression screening, and explainability-based fairness interventions often trade away accuracy without reliable fairness gains.

  2. Gender Fairness of Machine Learning Algorithms for Pain Detection

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Across four classifiers trained on the UNBC shoulder-pain dataset, every model showed gender disparities in pain detection, with the Vision Transformer achieving the best accuracy and some fairness metrics but not all.

Pith tools