REVIEW 8 cited by
An LLM Compiler for Parallel Function Calling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The reasoning capabilities of the recent LLMs enable them to execute external function calls to overcome their inherent limitations, such as knowledge cutoffs, poor arithmetic skills, or lack of access to private data. This development has allowed LLMs to select and coordinate multiple functions based on the context to tackle more complex problems. However, current methods for function calling often require sequential reasoning and acting for each function which can result in high latency, cost, and sometimes inaccurate behavior. To address this, we introduce LLMCompiler, which executes functions in parallel to efficiently orchestrate multiple function calls. Drawing inspiration from the principles of classical compilers, LLMCompiler enables parallel function calling with three components: (i) a Function Calling Planner, formulating execution plans for function calling; (ii) a Task Fetching Unit, dispatching function calling tasks; and (iii) an Executor, executing these tasks in parallel. LLMCompiler automatically generates an optimized orchestration for the function calls and can be used with both open-source and closed-source models. We have benchmarked LLMCompiler on a range of tasks with different patterns of function calling. We observe consistent latency speedup of up to 3.7x, cost savings of up to 6.7x, and accuracy improvement of up to ~9% compared to ReAct. Our code is available at https://github.com/SqueezeAILab/LLMCompiler.
Forward citations
Cited by 8 Pith papers
-
Auto: The AGI Compiler
AUTO compiles witnessed-deterministic LLM-agent spans into verified WASM cognition binaries and recompiles on deopt, cutting cost 6.4× at 96.9% parity on a 300-item shifted stream.
-
TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows
TraceCompiler recovers producer-consumer dependencies from noisy agent traces by admitting only uniquely-justified data flow and abstaining on ambiguity, achieving 0.928 precision on T1 versus 0.711 F1 for adjacency.
-
DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering
DrafterBench is a new benchmark of 1,920 PDF drawing-revision tasks; on it, the best model (OpenAI o1) averages about 80/100, and all tested models fail hard on incomplete instructions and plan execution.
-
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography
A three-step agentic workflow with LLM function calling and reflection improved glaucoma classification, CDR estimation, and repeatability over LLM-alone baselines, approaching specialist-level accuracy.
-
What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
The paper defines prompt graph engineering via four necessary and sufficient conditions (explicit structure, structure/content separation, executable semantics, first-class artifact) and an inclusion/exclusion test th...
-
TimelyLLM: Segmented LLM Serving System for Time-sensitive Robotic Applications
TimelyLLM segments LLM-generated robot plans into executable pieces and schedules those pieces by urgency, reducing response delays for time-critical robot tasks.
-
Less is More: Optimizing Function Calling for LLM Execution on Edge Devices
Selectively reducing the number of tools presented to an LLM, using embedding similarity over individual tools or clusters, improves function-calling success and efficiency on edge devices.
-
LLM Enabled Multi-Agent System for 6G Networks: Framework and Method of Dual-Loop Edge-Terminal Collaboration
A dual-loop edge-terminal multi-agent framework, combining task decomposition with parallel tool calling and offloading, is shown in a simulated 6G urban safety case study to outperform ReAct and LLMCompiler.
Discussion (0). Continue with ORCID to comment.