Pith. sign in

hub

ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

11 Pith papers cite this work, alongside 6 external citations. Polarity classification is still indexing.

11 Pith papers citing it
6 external citations · OpenAlex

hub tools

years

2026 11

representative citing papers

LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

cs.AI · 2026-06-18 · unverdicted · novelty 6.0

LedgerAgent is an inference-time method that uses a structured ledger to track task states and enforce domain policies in tool-calling agents, improving average pass^k over standard prompt-based approaches across four domains.

Scaling Agentic Capabilities via Grounded Interaction Synthesis

cs.CL · 2026-06-01 · unverdicted · novelty 6.0

GAIS synthesizes diverse, high-fidelity agentic tasks from real-world MCP servers and adversarial planning, outperforming LLM-only baselines on BFCL, τ²-Bench, and ACEBench with greater data efficiency.

Self-evolving LLM agents with in-distribution Optimization

cs.LG · 2026-06-05 · unverdicted · novelty 5.0

Q-Evolve unifies automatic process-reward labeling via advantage estimation and behavior-proximal policy optimization inside an in-distribution RL loop to enable self-evolving LLM agents on interactive tasks.

citing papers explorer

Showing 11 of 11 citing papers.