Autoresearch agent discovers objective-dependent LLM pipelines for sequential social dilemmas that exceed hand-designed baselines and inject fairness mechanisms only under maximin welfare.
Cooperation and Exploitation in
2 Pith papers cite this work. Polarity classification is still indexing.
abstract
We propose an LLM harness that generates code-based policy functions for multi-agent environments, evaluates them with self-play, and refines them using feedback from previous iterations. Following the recent line of work in feedback engineering (the design of which information signals are shown to the LLM during refinement), we compare sparse feedback (scalar reward only) with dense feedback (reward plus social metrics: efficiency, equality, sustainability, peace). In two Sequential Social Dilemmas (Gathering and Cleanup) and with two frontier LLMs (Claude Sonnet 4.6, Gemini 3.1 Pro), dense feedback improves over or matches sparse feedback on all metrics. We explain this asymmetry via feedback aliasing: when the scalar reward maps distinct failure modes into the same value (e.g., under- vs. over-cleaning), social metrics disambiguate and allow the LLM to diagnose which direction of improvement to take. We conclude that social metrics act as a coordination signal, leading to strategies such as Voronoi territory partitioning and adaptive cleaner schedules. Code at https://github.com/vicgalle/llm-policies-social-dilemmas.
citation-role summary
citation-polarity summary
years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
Presents Metal-Sci benchmark and harness for evolutionary LLM kernel search on Apple Silicon Metal, reporting in-distribution speedups up to 10.7x and using held-out gate scoring as oversight.
citing papers explorer
-
Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas
Autoresearch agent discovers objective-dependent LLM pipelines for sequential social dilemmas that exceed hand-designed baselines and inject fairness mechanisms only under maximin welfare.
-
Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon
Presents Metal-Sci benchmark and harness for evolutionary LLM kernel search on Apple Silicon Metal, reporting in-distribution speedups up to 10.7x and using held-out gate scoring as oversight.