Skill-Use is a 79-skill, 177-task benchmark showing that LLM agents fail to reliably retrieve, follow, and respect the boundaries of skills under progressive disclosure, with harness choice shifting model rankings.
Theagentcompany: benchmarking llm agents on consequen- tial real world tasks.Advances in Neural Information Processing Systems, 38, 2026
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?
Skill-Use is a 79-skill, 177-task benchmark showing that LLM agents fail to reliably retrieve, follow, and respect the boundaries of skills under progressive disclosure, with harness choice shifting model rankings.