Correlation analysis of four RAG metric libraries against human evaluations on 96 questions shows RAGChecker metrics correlate strongly with human scores, while generation-as-overall metrics correlate weakly, but single-system confounding limits interpretability.
Dynamic Knowledge Integration for Evidence-Driven Counter-Argument Generation with Large Language Models
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This paper investigates the role of dynamic external knowledge integration in improving counter-argument generation using Large Language Models (LLMs). While LLMs have shown promise in argumentative tasks, their tendency to generate lengthy, potentially unfactual responses highlights the need for more controlled and evidence-based approaches. We introduce a new manually curated dataset of argument and counter-argument pairs specifically designed to balance argumentative complexity with evaluative feasibility. We also propose a new LLM-as-a-Judge evaluation methodology that shows a stronger correlation with human judgments compared to traditional reference-based metrics. Our experimental results demonstrate that integrating dynamic external knowledge from the web significantly improves the quality of generated counter-arguments, particularly in terms of relatedness, persuasiveness, and factuality. The findings suggest that combining LLMs with real-time external knowledge retrieval offers a promising direction for developing more effective and reliable counter-argumentation systems.
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluating RAG Metrics in Applied Contexts: An Experiment, Its Findings and Its Limitations
Correlation analysis of four RAG metric libraries against human evaluations on 96 questions shows RAGChecker metrics correlate strongly with human scores, while generation-as-overall metrics correlate weakly, but single-system confounding limits interpretability.