REVIEW 7 cited by
TDAG: A Multi-Agent Framework based on Dynamic Task Decomposition and Agent Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The emergence of Large Language Models (LLMs) like ChatGPT has inspired the development of LLM-based agents capable of addressing complex, real-world tasks. However, these agents often struggle during task execution due to methodological constraints, such as error propagation and limited adaptability. To address this issue, we propose a multi-agent framework based on dynamic Task Decomposition and Agent Generation (TDAG). This framework dynamically decomposes complex tasks into smaller subtasks and assigns each to a specifically generated subagent, thereby enhancing adaptability in diverse and unpredictable real-world tasks. Simultaneously, existing benchmarks often lack the granularity needed to evaluate incremental progress in complex, multi-step tasks. In response, we introduce ItineraryBench in the context of travel planning, featuring interconnected, progressively complex tasks with a fine-grained evaluation system. ItineraryBench is designed to assess agents' abilities in memory, planning, and tool usage across tasks of varying complexity. Our experimental results reveal that TDAG significantly outperforms established baselines, showcasing its superior adaptability and context awareness in complex task scenarios.
Forward citations
Cited by 7 Pith papers
-
Single-agent or Multi-agent Systems? Why Not Both?
On 15 agentic benchmarks, the accuracy advantage of multi-agent LLM systems over single-agent systems mostly disappears with stronger base models, and a hybrid single/multi-agent cascade improves accuracy and cuts cost.
-
Adaptive Graph of Thoughts: Test-Time Adaptive Reasoning Unifying Chain, Tree, and Graph Structures
AGoT is a recursive graph-based prompting framework that decomposes LLM queries into nested subgraphs and reports large relative gains on some benchmarks, though headline GPQA gains rely on a shuffled subset.
-
COALESCE: Economic and Security Dynamics of Skill-Based Task Outsourcing Among Team of Autonomous LLM Agents
COALESCE, a framework for skill-based task outsourcing among LLM agents, claims 41.8% simulated and 20.3% real cost reductions, but the validation contains internal contradictions.
-
AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement
An LLM+knowledge-graph+human-in-the-loop framework improves simulated task-completion success over LLM-only and LLM+KG baselines, though the human oracle inflates the reported gains.
-
Flow: Modularized Agentic Workflow Automation
Flow represents a task as a dependency graph of subtasks and lets LLM agents redraw that graph during execution, reporting better success rates than three baselines on three coding tasks.
-
RedTeamLLM: an Agentic AI framework for offensive security
The paper reports that adding a separate reasoning step to a terminal-operating LLM agent reduces tool calls and improves completion on 4 of 5 entry-level CTF virtual machines, while the framework's memory and plan-co...
-
Adaptable and Precise: Enterprise-Scenario LLM Function-Calling Capability Training Pipeline
A LoRA fine-tuned 7B model trained on AI-synthesized enterprise HR API data beat GPT-4 and GPT-4o on the authors' private benchmark.
Discussion (0). Continue with ORCID to comment.