Pith. sign in

REVIEW 2 cited by

Robust Planning with Compound LLM Architectures: An LLM-Modulo Approach

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.14484 v1 pith:DHPJGU6S submitted 2024-11-20 cs.CL cs.AI

classification cs.CLcs.AI
keywords frameworkoutputcompoundllm-moduloperformanceworkapproacharchitectures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Previous work has attempted to boost Large Language Model (LLM) performance on planning and scheduling tasks through a variety of prompt engineering techniques. While these methods can work within the distributions tested, they are neither robust nor predictable. This limitation can be addressed through compound LLM architectures where LLMs work in conjunction with other components to ensure reliability. In this paper, we present a technical evaluation of a compound LLM architecture--the LLM-Modulo framework. In this framework, an LLM is paired with a complete set of sound verifiers that validate its output, re-prompting it if it fails. This approach ensures that the system can never output any fallacious output, and therefore that every output generated is guaranteed correct--something previous techniques have not been able to claim. Our results, evaluated across four scheduling domains, demonstrate significant performance gains with the LLM-Modulo framework using various models. Additionally, we explore modifications to the base configuration of the framework and assess their impact on overall system performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fitting the Message to the Moment: Designing Calendar-Aware Stress Messaging with Large Language Models

    cs.HC 2025-05 conditional novelty 5.0 of 10

    A one-week probe with eight students found that LLM-generated calendar-aware stress messages are valued when they target truly stressful events and use a concise, colloquial tone.

  2. Large Language Models for Planning: A Comprehensive and Systematic Survey

    cs.AI 2025-05 conditional novelty 3.0 of 10

    A structured survey of LLM planning methods, benchmarks, and interpretability work, organized around a three-way taxonomy.

Pith tools