Pith. sign in

REVIEW 2 cited by

Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.12952 v2 pith:WDHMVFMY submitted 2024-10-16 cs.CL

classification cs.CL
keywords taskscompositionalfunctionllmsatomiccallingfunctionsinstruction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have exhibited significant potential in performing diverse tasks, including the ability to call functions or use external tools to enhance their performance. While current research on function calling by LLMs primarily focuses on single-turn interactions, this paper addresses the overlooked necessity for LLMs to engage in multi-turn function calling--critical for handling compositional, real-world queries that require planning with functions but not only use functions. To facilitate this, we introduce an approach, BUTTON, which generates synthetic compositional instruction tuning data via bottom-up instruction construction and top-down trajectory generation. In the bottom-up phase, we generate simple atomic tasks based on real-world scenarios and build compositional tasks using heuristic strategies based on atomic tasks. Corresponding function definitions are then synthesized for these compositional tasks. The top-down phase features a multi-agent environment where interactions among simulated humans, assistants, and tools are utilized to gather multi-turn function calling trajectories. This approach ensures task compositionality and allows for effective function and trajectory generation by examining atomic tasks within compositional tasks. We produce a dataset BUTTONInstruct comprising 8k data points and demonstrate its effectiveness through extensive experiments across various LLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models

    cs.AI 2025-06 reject novelty 4.0 of 10

    KunLunBaizeRAG reports improved exact-match and LLM-judged scores on four multi-hop QA benchmarks using a reinforcement-learning-driven RAG framework.

  2. When Meaning Stays the Same, but Models Drift: Evaluating Quality of Service under Token-Level Behavioral Instability in LLMs

    cs.CL 2025-06 reject novelty 4.0 of 10

    LLM outputs drift measurably when prompts are reworded without changing meaning, instruction-tuned models drift less, and the new PBSS score quantifies this drift using embedding distance.

Pith tools