A single LoRA fine-tuned Llama-3.1-8B-Instruct model trained jointly on eight argument-mining tasks across 19 datasets matches or beats task-specific models, and merged models offer a cheaper compromise.
MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and Classification
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We introduce MAFALDA, a benchmark for fallacy classification that merges and unites previous fallacy datasets. It comes with a taxonomy that aligns, refines, and unifies existing classifications of fallacies. We further provide a manual annotation of a part of the dataset together with manual explanations for each annotation. We propose a new annotation scheme tailored for subjective NLP tasks, and a new evaluation method designed to handle subjectivity. We then evaluate several language models under a zero-shot learning setting and human performances on MAFALDA to assess their capability to detect and classify fallacies.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
AMELIA: A Family of Multi-task End-to-end Language Models for Argumentation
A single LoRA fine-tuned Llama-3.1-8B-Instruct model trained jointly on eight argument-mining tasks across 19 datasets matches or beats task-specific models, and merged models offer a cheaper compromise.