About 21% of OpenHands agent trajectories on SetupBench contained at least one insecure action, CWE-200 information exposure was the most common weakness, and GPT-4.1 showed the highest mitigation success at 96.8%.
https://windsurf.com
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
When Developer Aid Becomes Security Debt: A Systematic Analysis of Insecure Behaviors in LLM Coding Agents
About 21% of OpenHands agent trajectories on SetupBench contained at least one insecure action, CWE-200 information exposure was the most common weakness, and GPT-4.1 showed the highest mitigation success at 96.8%.