SoCRATES introduces a benchmark for proactive LLM mediators across eight domains and five socio-cognitive axes with topic-localized evaluation, finding top models close only about one-third of the unmediated consensus gap.
Robots in the middle: Evaluating llms in dispute resolution.arXiv preprint arXiv:2410.07053,
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
Humans reach 64.8% accuracy detecting synthetic legal evidence images overall but drop to chance levels on top generators, while MLLMs achieve 100% specificity yet only 5.9% detection on the hardest synthetics, with uncorrelated error patterns.
LLM facilitation in group charity allocation leaves consensus and participation equity unchanged while shifting specific allocations up to 5.5 points and increasing perceived trust.
citing papers explorer
-
SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations
SoCRATES introduces a benchmark for proactive LLM mediators across eight domains and five socio-cognitive axes with topic-localized evaluation, finding top models close only about one-third of the unmediated consensus gap.
-
Can You Trust What You See? Human and AI Detection of Synthetic Legal Evidence
Humans reach 64.8% accuracy detecting synthetic legal evidence images overall but drop to chance levels on top generators, while MLLMs achieve 100% specificity yet only 5.9% detection on the hardest synthetics, with uncorrelated error patterns.
-
Real-Time Group Dynamics with LLM Facilitation: Evidence from a Charity Allocation Task
LLM facilitation in group charity allocation leaves consensus and participation equity unchanged while shifting specific allocations up to 5.5 points and increasing perceived trust.