The paper introduces PolicySimEval, a claimed first benchmark with 20 comprehensive, 65 targeted, and 200 auto-generated tasks, and reports that GPT-4o and Llama-based agents score below 25% on most evaluation metrics.
Agent based modeling for agricultural policy evaluation: A review,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.MA 1years
2025 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
PolicySimEval: A Benchmark for Evaluating Policy Outcomes through Agent-Based Simulation
The paper introduces PolicySimEval, a claimed first benchmark with 20 comprehensive, 65 targeted, and 200 auto-generated tasks, and reports that GPT-4o and Llama-based agents score below 25% on most evaluation metrics.