GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

· 2026 · cs.AI · arXiv 2602.12316

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

open full Pith review browse 3 citing papers arXiv PDF

abstract

Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. Across 15 frontier models, agents fail to choose socially beneficial actions in 38% of high-stakes cases, such as military escalation, election manipulation, and medical malpractice. We measure sensitivity to game-theoretic prompt framing and ordering, and analyze reasoning patterns driving failures. We further show that game-theoretic interventions improve socially beneficial outcomes by up to 18%. Our results highlight substantial reliability gaps and provide a broad standardized testbed for studying alignment in multi-agent environments. The benchmark and code are available at https://github.com/causalNLP/gt-harmbench.

representative citing papers

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance

nlin.AO · 2026-05-17 · unverdicted · novelty 7.0

LLM societies in Nomic show non-monotonic collective adaptation peaking at mid-scales, with smaller models rule-inert and larger ones restrictive.

Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI

cs.GT · 2026-05-08 · unverdicted · novelty 6.0 · 2 refs

Mechanism design leaves a strictly positive welfare loss under incomplete contracts, but prosocial LLM agents close the gap in resource allocation and social dilemma settings.

Strategic Heterogeneous Multi-Agent Architecture for Cost-Effective Code Vulnerability Detection

cs.CR · 2026-04-23 · unverdicted · novelty 5.0

A game-theoretic heterogeneous multi-agent architecture with three cloud LLMs and a local verifier achieves 77.2% F1, 100% recall, and 3x speedup for code vulnerability detection at $0.002 per sample on the NIST Juliet suite.

citing papers explorer

Showing 3 of 3 citing papers after filters.

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance nlin.AO · 2026-05-17 · unverdicted · none · ref 15 · internal anchor
LLM societies in Nomic show non-monotonic collective adaptation peaking at mid-scales, with smaller models rule-inert and larger ones restrictive.
Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI cs.GT · 2026-05-08 · unverdicted · none · ref 2 · 2 links · internal anchor
Mechanism design leaves a strictly positive welfare loss under incomplete contracts, but prosocial LLM agents close the gap in resource allocation and social dilemma settings.
Strategic Heterogeneous Multi-Agent Architecture for Cost-Effective Code Vulnerability Detection cs.CR · 2026-04-23 · unverdicted · none · ref 1 · internal anchor
A game-theoretic heterogeneous multi-agent architecture with three cloud LLMs and a local verifier achieves 77.2% F1, 100% recall, and 3x speedup for code vulnerability detection at $0.002 per sample on the NIST Juliet suite.

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

fields

years

verdicts

representative citing papers

citing papers explorer