Pith. sign in

Theagentcompany: benchmarking llm agents on consequen- tial real world tasks.Advances in Neural Information Processing Systems, 38, 2026

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 1

years

2026 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

cs.CL · 2026-08-05 · conditional · novelty 6.0

Skill-Use is a 79-skill, 177-task benchmark showing that LLM agents fail to reliably retrieve, follow, and respect the boundaries of skills under progressive disclosure, with harness choice shifting model rankings.

citing papers explorer

Showing 1 of 1 citing paper.

  • Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses? cs.CL · 2026-08-05 · conditional · none · ref 2

    Skill-Use is a 79-skill, 177-task benchmark showing that LLM agents fail to reliably retrieve, follow, and respect the boundaries of skills under progressive disclosure, with harness choice shifting model rankings.