{"id":"0ce3d9f2-db5b-4d2a-a05e-b8aee1ea6580","arxiv_id":"2502.09100","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of logical reasoning in large language models that organizes benchmarks, evaluations, and enhancement methods around formal and symbolic logic.","lead":"This paper reviews how large language models handle logical reasoning, covering benchmarks, evaluation results, and methods to improve reasoning. It is useful as a structured map of the field for researchers and engineers entering this area.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim to be a comprehensive, formal-logic-focused review is not yet supported: no selection methodology is given, and several included 'logical reasoning' methods are general heuristics, so the survey's defining boundary is not operationalized.","rationale":"The paper is a useful and generally well-organized synthesis; it cites many relevant benchmarks and methods, and the neuro-symbolic taxonomy in §5.4 is a reasonable organizing device. I would not reject it. However, the survey's reason to exist is the distinction between logical reasoning and general heuristics. That distinction is asserted in §1 but never made operational: there is no inclusion/exclusion protocol, and §5.2 places methods on the heuristic side of its own divide (e.g., Maieutic Prompting, Logic-of-Thoughts, GoT) under the logical-reasoning umbrella. The reader's concern about undocumented selection is real, but my emphasis is the internal boundary: even if the literature search were perfect, the paper would not have a principled answer to which papers belong in a formal-logic survey. Table 1's source-label errors (ProofWriter and LogicNLI) reinforce the impression that the corpus was not systematically audited. These problems are addressable through explicit criteria and re-verification, so the existing CONDITIONAL verdict remains appropriate; no verdict change is needed.","tokens_in":15052,"tokens_out":5665,"duration_ms":57489,"concrete_test":"Define a formal/symbolic criterion from the paper's own §1 framing: a method is in scope only if it translates to or invokes a formal language, theorem prover, rule engine, or explicit logical calculus. Apply this criterion to every method in §5.2's Inference-Time Decoding and §5.4. If Maieutic Prompting, Logic-of-Thoughts, GoT, and DetermLR fail the criterion, the paper must either exclude them or explicitly broaden its definition; if they pass, the authors should state the criterion under which they pass and apply it consistently to justify the central claim of focusing on formal and symbolic logic-based reasoning rather than general heuristic approaches.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is stated in §1: prior surveys conflate logical reasoning with general-purpose heuristic strategies, and this survey fills the gap by focusing on formal and symbolic logic-based reasoning rather than general heuristic approaches. For that claim to hold, the survey's inclusion boundary must be principled and its corpus representative. Neither is verifiable from the text. No systematic search, inclusion criteria, or audit trail is reported; the comprehensive label rests on an implicit selection. Table 1 compounds this: ProofWriter and LogicNLI are both labeled Exam-based despite §3.1 describing them as generated using logical rules, showing that the dataset inventory was not systematically verified. The boundary problem appears in §5.2: Maieutic Prompting, Logic-of-Thoughts, GoT, and DetermLR are placed under inference-time decoding as logical-reasoning enhancements, yet they do not use formal languages, theorem provers, or explicit logical calculi in the way that LINC or Logic-LM do. If these methods are in scope, the stated exclusion of general heuristics is not maintained; if they are out of scope, the survey has not justified their inclusion. Either way, the distinctiveness of the survey's contribution is weakened.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a survey of logical reasoning in large language models (LLMs), explicitly targeting formal and symbolic logic-based reasoning rather than general heuristic strategies such as chain-of-thought. The survey organizes the field into four reasoning paradigms (deductive, inductive, abductive, analogical), reviews existing benchmarks and datasets, summarizes evaluations of LLM capabilities across these paradigms, and surveys enhancement methods including data-centric tuning, reinforcement learning, inference-time decoding, external knowledge utilization, and neuro-symbolic approaches. It concludes with a discussion of open challenges and future directions. The central claim is that it fills a gap left by prior surveys that conflate logical reasoning with general-purpose heuristics.","tokens_in":15277,"tokens_out":3486,"duration_ms":34836,"significance":"If the survey's scope were carefully enforced, it would be a useful reference for researchers working at the intersection of formal logic and LLMs. It consolidates a dispersed literature, offers a structured taxonomy, and points to important benchmarks (FOLIO, ProofWriter, LogiQA, RuleTaker) and neuro-symbolic methods (LINC, Logic-LM, SymbCoT). The paper does not present new experimental results, but it attempts to synthesize and categorize a fast-moving area. The practical value, however, depends on whether the inclusion boundary between 'logical reasoning' and 'general heuristics' is defensible and whether the factual inventory in Table 1 is reliable. Those conditions are not currently met.","major_comments":[{"comment":"The survey's stated scope is formal and symbolic logic-based reasoning rather than general heuristic approaches, but §5.2 (Inference-Time Decoding) lists methods that do not use formal languages, theorem provers, or explicit logical calculi. In particular, Maieutic Prompting, Graph of Thought, Selection-Inference, and DetermLR are general inference-time scaffolding or heuristic methods, not symbolic-logic methods in the sense of LINC or Logic-LM. If these methods are considered logical-reasoning enhancements, the survey must explicitly justify why they fall under the formal/symbolic boundary; if they are not, their inclusion weakens the claim of a distinct focus. This is load-bearing because the survey's stated contribution is precisely to separate logical reasoning from general heuristics.","section":"§1 and §5.2"},{"comment":"Table 1 contains factual inconsistencies that undermine the reliability of the benchmark inventory. ProofWriter is labeled 'Exam-based' with size '—', but §3.1 describes it as extending RuleTaker with closed-world and open-world assumptions, i.e., a rule-generated dataset. LogicNLI is labeled 'Exam-based', yet §3.1 states it 'contains 30K entries generated using logical rules.' LogicBench is labeled 'Rule-based', but §3.1 says it is 'GPT-3-generated.' Additionally, GSM is listed as 'Exam-based' with 19K entries, but GSM8K is a collection of grade-school math word problems and GSM-PLUS is a robustness perturbation set; neither is a logical reasoning exam. These errors require correction and verification of all entries in the table.","section":"Table 1 and §3.1"},{"comment":"The paper claims to be 'a comprehensive review' but provides no systematic literature search strategy, inclusion criteria, or audit trail. Without such a methodology, the representativeness of the surveyed corpus cannot be assessed. This concern is exacerbated by the prominence of the authors' own prior work (ConTRoL, LogiQA, GLoRE, LogiCoT, Logic Agent) in the benchmark and enhancement sections, which raises the risk that the selection is shaped by the authors' research interests rather than by an objective coverage criterion. A short methodology subsection or appendix stating the search dates, databases, and inclusion/exclusion rules is needed to substantiate the comprehensiveness claim.","section":"§1 (methodology)"},{"comment":"The statement that DeepSeek-R1 'represents a potential paradigm shift in logical reasoning optimization' is an evaluative claim that is not supported by evidence presented in the survey. DeepSeek-R1 shows strong performance on general reasoning benchmarks, but the paper provides no demonstration that it specifically advances formal or symbolic logical reasoning (e.g., on FOLIO, ProofWriter, or similar logic-focused benchmarks). The claim should be tempered to describe observed performance gains and to note that questions about logical validity and consistency remain open.","section":"§5.2 (Reinforcement Learning)"}],"minor_comments":[{"comment":"Figure 1 contains a typo: 'RulteTaker' should be 'RuleTaker'; also, capitalization is inconsistent for 'Deepseek-R1-Zero' versus 'DeepSeek-R1'. The figure would benefit from a proofread pass.","section":"Figure 1"},{"comment":"Reference [liu et al., 2023d] is missing capitalization ('Liu et al.'); check that all citation keys and reference entries are spelled consistently, including 'Logic-of-Thoughts' versus 'Logic-of-thoughts' in the text.","section":"References"},{"comment":"The paragraph on 'GSM' is vague; it should clearly state that it covers GSM8K and GSM-PLUS and explain why these math word problems are categorized under logical reasoning, since they are not typically framed as logic benchmarks.","section":"§3.2"},{"comment":"In the Discussion, the sentence 'Data-driven methods (e.g., instruction tuning on LogicBench [Parmar et al., 2024])' cites the 2024 version of LogicBench, while the dataset is primarily introduced in the 2023 reference [Parmar et al., 2023]; please clarify the relationship between the two references.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important and timely topic, and the overall organization is reasonable. The main risk is not the absence of new results but the unverified factual content and the lack of a transparent selection protocol. The authors should treat the Table 1 corrections and the scope justifications as mandatory before the survey can be relied upon as a reference. I do not see a fatal flaw, but the current form is not yet suitable for publication in a serious journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Nice to see a survey that actually tries to separate formal/symbolic logic from the CoT-and-friends sprawl. That is the paper's real contribution, and the taxonomy (deductive/inductive/abductive/analogical) plus the benchmark and enhancement sections make it a plausible reference for newcomers. I agree with the reader's CONDITIONAL verdict. The internal organization is clear, and most claims are attributed.\n\nThe main soft spot is the inclusion boundary. The intro promises a focus on formal and symbolic logic rather than general heuristics, but §5.2 puts Maieutic Prompting, Graph-of-Thought, Logic-of-Thoughts, and DetermLR under inference-time decoding as logical reasoning enhancements. Those methods don't use formal languages, provers, or explicit calculi in the way LINC or Logic-LM do. Either the survey's definition is broader than advertised, or the selection is inconsistent. Either way, the distinctive claim weakens. A short operational definition of 'logical reasoning enhancement' with inclusion criteria would fix it.\n\nSecond, Table 1 has several factual errors. ProofWriter and LogicNLI are labeled Exam-based even though the text describes them as rule-generated; LogicBench says Rule-based but the text says GPT-3 generated; GSM is labeled Exam-based which is a stretch. These are exactly the kinds of errors that make a survey less trustworthy, and they are easy to fix. The missing ProofWriter size should also be filled.\n\nThird, the 'comprehensive' label is not supported by a documented search or audit trail. The heavy presence of the authors' own benchmarks (ConTRoL, LogiQA, GLoRE, LogiCoT, Logic Agent) may be justified, but without a methodology it's hard to know whether the corpus is representative. This is a minor-to-moderate issue; it doesn't invalidate the survey, but it does cap its authority.\n\nThe stress-test note is on point. The DeepSeek-R1 discussion drifts into hype ('potential paradigm shift'), but that's a rhetorical soft spot, not a load-bearing flaw.\n\nWho's this for? A grad student or researcher wanting an entry point to formal-logic-focused LLM reasoning work will get value. As a peer reviewer I'd ask for a definition of the scope, a corrected Table 1, and a short note on how papers were selected. The book is fixable. I'd accept it for review.","headline":"Useful survey with a real formal-logic angle, but the inclusion boundary and Table 1 need work before it can claim comprehensiveness.","tokens_in":15792,"tokens_out":2092,"would_cite":true,"duration_ms":20050,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that logical reasoning in large language models is a distinct field from chain-of-thought heuristics and maps it through four paradigms: deductive, inductive, abductive, and analogical.","keywords":["logical reasoning","large language models","deductive reasoning","inductive reasoning","abductive reasoning","analogical reasoning","neuro-symbolic reasoning","benchmarks"],"falsifier":"A systematic literature search with clear inclusion criteria that locates a substantial cluster of LLM logical-reasoning work outside the four paradigms—say modal, deontic, or temporal logic—or that shows researchers routinely treat chain-of-thought as logical reasoning would undercut the survey's central claim that a logic-focused review fills a real gap.","tokens_in":14860,"feed_emoji":"🧠","tokens_out":8114,"duration_ms":73843,"temperature":0.7,"pith_summary":"This survey works to establish that logical reasoning in large language models deserves treatment as its own research area, defined by formal and symbolic logic rather than by heuristic prompting techniques such as chain-of-thought. It organizes the field into four reasoning paradigms—deductive, inductive, abductive, and analogical—and sorts benchmarks and improvement methods into categories. The result is a map of what has been tested, where models succeed and fail, and which strategies (data tuning, reinforcement learning, decoding methods, and neuro-symbolic integration) are supposed to help. The survey concludes that heuristic performance is often strong but rigorous logical inference remains unreliable, especially under rephrasing and out-of-distribution inputs.","feed_headline":"Survey maps LLM logic through four reasoning types","feed_subtitle":"Deduction, induction, abduction, and analogy: where models pass, fail, and how training and solvers help.","key_machinery":"The organizing device is a two-dimensional map: the four reasoning paradigms (deductive, inductive, abductive, analogical) crossed with a benchmark typology (rule-based, expert-designed, exam-based) and an enhancement-method typology (data-centric, model-centric, external-knowledge, neuro-symbolic). The paper also formalizes each enhancement family as an optimization objective: data-centric tuning as $D^* = \\arg\\max_D R(M_D)$, model-centric optimization as $(\\theta^*, S^*) = \\arg\\max_{\\theta,S} R(M_\\theta, S)$, and neuro-symbolic coupling as $(M^*, P^*)$ in which the LLM maps natural language to a formal language $z = M(x)$ and the solver produces $y = P(z)$. These equations make explicit what each strategy is supposed to optimize, and the taxonomy is what organizes the survey's evidence.","core_discovery":"The paper's central claim is that logical reasoning in LLMs is a distinct object of study, separate from general reasoning surveys, because the field's core is formal and symbolic inference rather than plausible text generation. On its reading, the state of the art is mixed: models score well on many benchmarks across all four paradigms, but fail on rephrased, out-of-distribution, or extended reasoning tasks, indicating reliance on surface-level statistical patterns. The survey argues that progress comes from three complementary levers: better data (expert-curated, synthetic, and LLM-distilled), better models (instruction tuning and reinforcement learning), and better inference-time machinery (constrained decoding, external solvers, and neuro-symbolic pipelines). It identifies evaluation as the open frontier, since current accuracy-based metrics conflate reasoning with pattern recognition and do not test consistency, soundness, or generalization.","pith_inferences":["Editorial extension: because the survey's evidence repeatedly ties reasoning failures to paraphrase sensitivity, one testable prediction follows—models trained on logically annotated data should survive rephrasing better than models trained on chain-of-thought traces alone.","Editorial extension: the four-paradigm taxonomy omits modal, deontic, temporal, and defeasible logics; checking whether recent LLM work in those areas fits the map would test the taxonomy's completeness.","Editorial extension: the survey draws often on benchmarks and systems developed by its own authors, so an independent meta-analysis with transparent inclusion criteria would be the natural next verification of the field map."],"forward_implications":["Benchmark design should separate logical competence from pattern recognition by adding perturbed rephrasings, negated premises, and swapped quantifiers to evaluation suites.","Improvement methods should be compared across tasks rather than on single benchmarks, since the survey's evidence shows that gains from instruction tuning and reinforcement learning are often task-bound.","Neuro-symbolic pipelines that translate natural language into formal logic and hand off to solvers or rule-guided LLM chains offer a route to verifiable reasoning, not just higher multiple-choice scores.","Evaluation metrics should move from accuracy alone toward consistency and soundness, because current multiple-choice formats conflate reasoning ability with memorization."],"supporting_citations":[{"why":"Cited as an existing reasoning survey that treats reasoning broadly, motivating the paper's narrower logic-focused scope.","marker":"[Plaat et al., 2024]"},{"why":"Another foundation-model reasoning survey that the paper distinguishes from a logic-specific review.","marker":"[Sun et al., 2023]"},{"why":"Supplies the NeuLR evaluation framework and a six-dimension assessment of LLM logical reasoning.","marker":"[Xu et al., 2023]"},{"why":"Provides evidence on deductive reasoning limitations of LLMs on out-of-distribution examples.","marker":"[Saparov et al., 2023]"},{"why":"FOLIO, the expert-constructed first-order logic benchmark, anchors formal-logic evaluation and data-centric training.","marker":"[Han et al., 2024a]"},{"why":"Logic-LM, a reference neuro-symbolic pipeline that couples LLMs with symbolic solvers.","marker":"[Pan et al., 2023]"},{"why":"LINC, an early neuro-symbolic method converting natural language to first-order logic for theorem proving.","marker":"[Olausson et al., 2023]"},{"why":"The reinforcement-learning reasoning system used to frame the current frontier of reasoning-oriented LLMs.","marker":"[DeepSeek-AI, 2025]"}],"fun_headline_variants":["LLM logic: strong on benchmarks, weak on rephrases, survey says","Four reasoning paradigms, one survey: where LLMs fall short","Logical reasoning in LLMs: benchmarks fool us, survey warns","Why LLMs ace logic tests but trip on twists: survey explains"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The map is only as trustworthy as the set of papers the authors chose to include, and the survey reports no systematic search or inclusion criteria that would prove the four-paradigm taxonomy covers the field.","fun_headline_variants_meta":{"raw":{"variants":["LLM logic: strong on benchmarks, weak on rephrases, survey says","Four reasoning paradigms, one survey: where LLMs fall short","Logical reasoning in LLMs: benchmarks fool us, survey warns","Why LLMs ace logic tests but trip on twists: survey explains"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000709,"raw_usage":{"total_tokens":3140,"prompt_tokens":842,"completion_tokens":2298,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":458,"completion_tokens_details":{"reasoning_tokens":2222}},"tokens_in":458,"tokens_out":2298,"duration_ms":14855,"temperature":1.0,"reasoning_tokens":2222,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T22:35:53.602364+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic literature search with clear inclusion criteria that locates a substantial cluster of LLM logical-reasoning work outside the four paradigms—say modal, deontic, or temporal logic—or that shows researchers routinely treat chain-of-thought as logical reasoning would undercut the survey's central claim that a logic-focused review fills a real gap.","supporting_citations":[{"cited_title":"Testing the general de- ductive reasoning capacity of large language models using ood examples","cited_arxiv_id":null,"evidence_quote":"Provides evidence on deductive reasoning limitations of LLMs on out-of-distribution examples."},{"cited_title":"Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning","cited_arxiv_id":null,"evidence_quote":"Logic-LM, a reference neuro-symbolic pipeline that couples LLMs with symbolic solvers."},{"cited_title":"LINC: A neurosymbolic approach for logical reasoning by combining language models with first-order logic provers","cited_arxiv_id":null,"evidence_quote":"LINC, an early neuro-symbolic method converting natural language to first-order logic for theorem proving."},{"cited_title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","cited_arxiv_id":null,"evidence_quote":"The reinforcement-learning reasoning system used to frame the current frontier of reasoning-oriented LLMs."}],"review_version":1}