A new expert-annotated dataset with span-level depression labels is used to compare GPT-4.1, Claude 3.7, and Gemini 2.5 Pro on explanation faithfulness, finding no consistent gain from few-shot prompting.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Gold Standard Dataset and Evaluation Framework for Depression Detection and Explanation in Social Media using LLMs
A new expert-annotated dataset with span-level depression labels is used to compare GPT-4.1, Claude 3.7, and Gemini 2.5 Pro on explanation faithfulness, finding no consistent gain from few-shot prompting.