Pith. sign in

StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The rapid advancement of large language models (LLMs) has spurred significant interest in tool learning, where LLMs are augmented with external tools to tackle complex tasks. However, existing tool environments face challenges in balancing stability, scalability, and realness, particularly for benchmarking purposes. To address this problem, we propose MirrorAPI, a novel framework that trains specialized LLMs to accurately simulate real API responses, effectively acting as "mirrors" to tool environments. Using a comprehensive dataset of request-response pairs from 7,000+ APIs, we employ supervised fine-tuning and chain-of-thought reasoning to enhance simulation fidelity. MirrorAPI achieves superior accuracy and stability compared to state-of-the-art methods, as demonstrated by its performance on the newly constructed MirrorAPI-Bench and its integration into StableToolBench.

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

TextAtari: 100K Frames Game Playing with Language Agents

cs.CL · 2025-06-04 · conditional · novelty 5.0

TextAtari is a text-based Atari benchmark for language agents; 7-8B LLMs stay below 10% of human scores in over 90% of tested conditions, and knowledge injection helps more than chain-of-thought.

citing papers explorer

Showing 1 of 1 citing paper.

  • TextAtari: 100K Frames Game Playing with Language Agents cs.CL · 2025-06-04 · conditional · none · ref 44 · internal anchor

    TextAtari is a text-based Atari benchmark for language agents; 7-8B LLMs stay below 10% of human scores in over 90% of tested conditions, and knowledge injection helps more than chain-of-thought.