Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.AI 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Interpretable Risk Mitigation in LLM Agent Systems

cs.AI · 2025-05-15 · conditional · novelty 6.0

Steering LLaMA-3-8B's internal 'good faith/bad faith' feature shifts its defection probability in the iterated prisoner's dilemma by 28 percentage points.

citing papers explorer

Showing 1 of 1 citing paper.

  • Interpretable Risk Mitigation in LLM Agent Systems cs.AI · 2025-05-15 · conditional · none · ref 6

    Steering LLaMA-3-8B's internal 'good faith/bad faith' feature shifts its defection probability in the iterated prisoner's dilemma by 28 percentage points.