Pith. sign in

REVIEW 2 cited by

Boosting Logical Reasoning in Large Language Models through a New Framework: The Graph of Thought

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.08614 v1 pith:2AJ5JVF6 submitted 2023-08-16 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords accuracymodelspromptingtextitgpt-4graphlogicalmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Recent advancements in large-scale models, such as GPT-4, have showcased remarkable capabilities in addressing standard queries. However, when facing complex problems that require multi-step logical reasoning, their accuracy dramatically decreases. Current research has explored the realm of \textit{prompting engineering} to bolster the inferential capacities of these models. Our paper unveils a pioneering prompting technique, dubbed \textit{Graph of Thoughts (GoT)}. Through testing on a trio of escalating challenges: the 24-point game, resolution of high-degree polynomial equations, and derivation of formulas for recursive sequences, our method outperformed GPT-4, achieving accuracy improvements of $89.7\%$, $86\%$, and $56\%$ for each respective task. Moreover, when juxtaposed with the state-of-the-art (SOTA) prompting method, \textit{Tree of Thought (ToT)}, our approach registered an average accuracy boost of $23\%$, $24\%$, and $15\%$.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs

    cs.CL 2026-07 conditional novelty 6.5 of 10

    On 2,520 programming tasks, matched Qwen general and coder models reliably raise Bloom cognitive demand but fail to lower it, so execution skill does not imply educational control.

  2. DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching

    cs.SD 2025-08 unverdicted novelty 4.0 of 10

    DAFMSVC swaps source SSL features for similar target features and adds dual cross-attention plus flow matching to improve one-shot singing voice conversion.

Pith tools