Pith. sign in

REVIEW 7 cited by

TDAG: A Multi-Agent Framework based on Dynamic Task Decomposition and Agent Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10178 v2 pith:BWAHI7VJ submitted 2024-02-15 cs.CL

classification cs.CL
keywords taskscomplextaskadaptabilityagentsframeworktdagagent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The emergence of Large Language Models (LLMs) like ChatGPT has inspired the development of LLM-based agents capable of addressing complex, real-world tasks. However, these agents often struggle during task execution due to methodological constraints, such as error propagation and limited adaptability. To address this issue, we propose a multi-agent framework based on dynamic Task Decomposition and Agent Generation (TDAG). This framework dynamically decomposes complex tasks into smaller subtasks and assigns each to a specifically generated subagent, thereby enhancing adaptability in diverse and unpredictable real-world tasks. Simultaneously, existing benchmarks often lack the granularity needed to evaluate incremental progress in complex, multi-step tasks. In response, we introduce ItineraryBench in the context of travel planning, featuring interconnected, progressively complex tasks with a fine-grained evaluation system. ItineraryBench is designed to assess agents' abilities in memory, planning, and tool usage across tasks of varying complexity. Our experimental results reveal that TDAG significantly outperforms established baselines, showcasing its superior adaptability and context awareness in complex task scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Single-agent or Multi-agent Systems? Why Not Both?

    cs.MA 2025-05 conditional novelty 6.0 of 10

    On 15 agentic benchmarks, the accuracy advantage of multi-agent LLM systems over single-agent systems mostly disappears with stronger base models, and a hybrid single/multi-agent cascade improves accuracy and cuts cost.

  2. Adaptive Graph of Thoughts: Test-Time Adaptive Reasoning Unifying Chain, Tree, and Graph Structures

    cs.AI 2025-02 conditional novelty 5.0 of 10

    AGoT is a recursive graph-based prompting framework that decomposes LLM queries into nested subgraphs and reports large relative gains on some benchmarks, though headline GPQA gains rely on a shuffled subset.

  3. COALESCE: Economic and Security Dynamics of Skill-Based Task Outsourcing Among Team of Autonomous LLM Agents

    cs.AI 2025-06 reject novelty 4.0 of 10

    COALESCE, a framework for skill-based task outsourcing among LLM agents, claims 41.8% simulated and 20.3% real cost reductions, but the validation contains internal contradictions.

  4. AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement

    cs.RO 2025-02 conditional novelty 4.0 of 10

    An LLM+knowledge-graph+human-in-the-loop framework improves simulated task-completion success over LLM-only and LLM+KG baselines, though the human oracle inflates the reported gains.

  5. Flow: Modularized Agentic Workflow Automation

    cs.AI 2025-01 conditional novelty 4.0 of 10

    Flow represents a task as a dependency graph of subtasks and lets LLM agents redraw that graph during execution, reporting better success rates than three baselines on three coding tasks.

  6. RedTeamLLM: an Agentic AI framework for offensive security

    cs.CR 2025-05 conditional novelty 3.0 of 10

    The paper reports that adding a separate reasoning step to a terminal-operating LLM agent reduces tool calls and improves completion on 4 of 5 entry-level CTF virtual machines, while the framework's memory and plan-co...

  7. Adaptable and Precise: Enterprise-Scenario LLM Function-Calling Capability Training Pipeline

    cs.AI 2024-12 reject novelty 3.0 of 10

    A LoRA fine-tuned 7B model trained on AI-synthesized enterprise HR API data beat GPT-4 and GPT-4o on the authors' private benchmark.

Pith tools