With a large DeBERTa model, annotating 500 to 1000 sentences per legal concept matches full annotation, and LLM-based annotation (Qwen 2.5) achieves NDCG scores close to or better than human-annotation-trained models on the statutory interpretation retrieval task.
Can GPT-4 Support Analysis of Textual Data in Tasks Requiring Highly Specialized Domain Expertise?
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We evaluated the capability of generative pre-trained transformers~(GPT-4) in analysis of textual data in tasks that require highly specialized domain expertise. Specifically, we focused on the task of analyzing court opinions to interpret legal concepts. We found that GPT-4, prompted with annotation guidelines, performs on par with well-trained law student annotators. We observed that, with a relatively minor decrease in performance, GPT-4 can perform batch predictions leading to significant cost reductions. However, employing chain-of-thought prompting did not lead to noticeably improved performance on this task. Further, we demonstrated how to analyze GPT-4's predictions to identify and mitigate deficiencies in annotation guidelines, and subsequently improve the performance of the model. Finally, we observed that the model is quite brittle, as small formatting related changes in the prompt had a high impact on the predictions. These findings can be leveraged by researchers and practitioners who engage in semantic/pragmatic annotations of texts in the context of the tasks requiring highly specialized domain expertise.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Are manual annotations necessary for statutory interpretations retrieval?
With a large DeBERTa model, annotating 500 to 1000 sentences per legal concept matches full annotation, and LLM-based annotation (Qwen 2.5) achieves NDCG scores close to or better than human-annotation-trained models on the statutory interpretation retrieval task.