Single-agent systems with tools provide the optimal performance-efficiency trade-off for small language models, outperforming base models and multi-agent setups.
Small LLM s Are Weak Tool Learners: A Multi- LLM Agent
4 Pith papers cite this work, alongside 19 external citations. Polarity classification is still indexing.
representative citing papers
PV-SQL boosts Text-to-SQL execution accuracy by 5% and valid efficiency by 20.8% on BIRD benchmarks via database probing and rule-based SQL verification while using fewer tokens.
A draft-and-repair agentic loop using GPT-4o-mini reduces text-to-chart execution errors to 4.5-4.6% on two benchmarks, suggesting execution is nearly solved and future work should focus on quality and accessibility.
A shuffle-based token classifier plus group-level loss reweighting improves supervised fine-tuning of LLM agents on tool-use benchmarks.
citing papers explorer
-
Rethinking Scale: Deployment Trade-offs of Small Language Models under Agent Paradigms
Single-agent systems with tools provide the optimal performance-efficiency trade-off for small language models, outperforming base models and multi-agent setups.
-
Does It Run and Is That Enough? Revisiting Text-to-Chart Generation with a Multi-Agent Approach
A draft-and-repair agentic loop using GPT-4o-mini reduces text-to-chart execution errors to 4.5-4.6% on two benchmarks, suggesting execution is nearly solved and future work should focus on quality and accessibility.
-
Disentangling Reasoning Tokens and Boilerplate Tokens For Language Model Fine-tuning
A shuffle-based token classifier plus group-level loss reweighting improves supervised fine-tuning of LLM agents on tool-use benchmarks.