REVIEW 4 major objections 5 minor 94 references
SymStep: Symbolic Step Verification for Logical Reasoning
T0 review · 4 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Per-step symbolic verification lets an LLM solve constraint-dense logic puzzles that defeat chain-of-thought, reaching 97% where direct prompting and CoT both score 0%.
desk verdict Clean, honest paper on a new per-step LLM + constraint-propagator loop with MRV guidance; strong headline numbers rest on small, filtered benchmarks, and the propagator checks consistency, not clue-entailment. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The DEDUCE protocol paired with a constraint propagator. The LLM produces a single atomic assignment per turn (e.g., 'DEDUCE: Alice, pet, Cat' or a negation), which is checked against a possibility matrix Π(p,a) of values still consistent with prior accepted deductions. The propagator applies positive and negative updates, rejects any update that empties a domain, and runs an arc-consistency cascade that automatically derives every forced assignment; SymStep+G then emits an MRV hint identifying the unresolved (entity, attribute) cell with the smallest candidate set. This machinery turns the LLM into a branching oracle within a deterministic constraint-satisfaction loop, outsourcing the bookk
What would settle it
Take any LGP-14 puzzle and, at some step, inject a plausible but clue-violating claim (e.g., 'DEDUCE: Alice, pet, Fish' when a clue forces Alice's pet elsewhere) while the propagator state still allows it; if the pipeline accepts it and later concludes with a unique but incorrect assignment passed off as verified, the soundness gap is confirmed.
Extended reading notes
Core claim
Under SymStep, reasoning is reorganized as a sequence of small checkable claims rather than a monologue. Each DEDUCE assertion is parsed and applied to a possibility matrix; the propagator eliminates values from other domains and, whenever a cell shrinks to one option, propagates that forced assignment without an extra LLM call. The paper's central empirical finding is that adding an MRV hint — which variable has the fewest remaining candidates — is what turns a reliable but deadlock-prone system (86% on LGP-14) into a perfect one (100%), and that the same pattern transfers to an external benchmark (0% to 97% on ZebraLogicBench). The authors interpret this as evidence that the bottleneck is
Load-bearing premise
The load-bearing premise is that the LLM's DEDUCE assertions are grounded in the puzzle clues; the propagator only checks consistency with previously accepted deductions, not entailment, so a plausible-but-false claim that happens to be consistent can still be accepted and propagated.
Editorial extensions
If this is right
- On constraint-dense benchmarks (LGP-14, SP-6, MWP-8, FIN-6, ZebraLogicBench, AR-LSAT), SymStep variants match or exceed every baseline; on algebra word problems they only match CoT, so the benefit is specific to constraint density.
- Ablations show MRV guidance alone reaches 100% with zero variance on LGP-6 while verification alone reaches 72%, making guidance the primary driver and the verifier a safety net that matters more as puzzles scale.
- The arc-consistency cascade derives about 60% of assignments automatically, so SymStep+G costs only about 7 LLM calls per puzzle (under two cents per puzzle on the reported setup) rather than one call per step.
- Whole-problem translation baselines (Logic-LM) score 0% on LGP-14 and SP-6, suggesting per-step interleaving with deterministic feedback is more robust than one-shot formalization.
- The propagator provides a formal consistency guarantee: no fact inconsistent with prior accepted claims ever enters the model's context.
Reading between the lines
- Inference: The same DEDUCE-plus-propagator loop should extend to temporal, counting, and conditional constraints by swapping in a SAT/ASP backend; the paper notes this as future work, but the protocol's design makes it a natural next step.
- Inference: The ablation result suggests that many LLM failures on constraint puzzles are search-direction failures, so simpler interventions — such as a fixed variable-ordering prompt that mimics MRV — might recover a large share of the gain without any symbolic machinery.
- Inference: The soundness gap (consistency vs. entailment) implies a testable upgrade: add a second check that verifies each accepted claim against the original clues, which would close the residual failure cases the paper attributes to consistent-but-wrong guesses.
- Inference: Because contradictions were near-zero under guidance, measuring the upper bound of puzzle size or clue density where guidance alone still yields 100% would be a clear stress test of where the safety net becomes essential.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SymStep, a neuro-symbolic protocol in which an LLM emits one atomic DEDUCE assertion per turn and a lightweight constraint propagator checks each assertion for consistency with previously accepted assertions, enforces arc-consistency, and cascades implied eliminations. SymStep+G augments this with an MRV (minimum remaining values) hint that tells the LLM which unresolved variable to tackle next. The authors evaluate the approach on a new hand-built LGP-14 benchmark, a 35-puzzle retained subset of ZebraLogicBench, AR-LSAT problems, and smaller scheduling, arithmetic, and financial benchmarks. Headline results include 100% on LGP-14 and 97% on the ZebraLogicBench subset versus 0% for Direct and CoT, and 100% vs. 87% for CoT on AR-LSAT. An ablation on LGP-6 attributes most of the gain to MRV guidance rather than contradiction feedback.
Significance. If the empirical results hold, SymStep is a simple, low-cost, and easy-to-reproduce method for substantially improving LLM performance on constraint-dense puzzles, a setting where CoT and whole-problem symbolic translation (Logic-LM) fail badly. The paper deserves credit for releasing code and a new benchmark (LGP-20), for reporting Wilson confidence intervals on small samples, for an N=3 multi-run ablation, and for including a domain-boundary experiment on AQUA-RAT that shows where the method does and does not help. The formal contribution is modest, however: the theorems establish only internal consistency and arc-consistency propagation, not that accepted deductions are entailed by the puzzle clues. The main intellectual interest is the interaction between a deterministic bookkeeping propagator and an LLM as a natural-language branching oracle.
major comments (4)
- [§3.2, Appendix F] The paper's central claim that SymStep 'verifies' each step is overstated. As the authors explicitly acknowledge in §3.2 and in the 'Soundness caveat' of Appendix F, the propagator checks consistency with previously accepted deductions, not logical entailment from the puzzle clues. A fabricated claim that is locally consistent but contradicts an unused clue is accepted, propagated, and used to compute MRV hints. Thus the method provides no guarantee that the final answer is correct; the observed 97-100% accuracies are properties of the LLM's guessing behavior on these distributions, not of the verification mechanism. The abstract and contributions use 'verifies'/'verified', and Appendix A claims 'deterministic soundness' — these need to be replaced with 'consistency checking' or the method must be extended to check clue-grounding. This is load-bearing because it determines what the paper
- [§4.9, abstract] The ZebraLogicBench result is based on a 'retained subset' of 35 puzzles obtained by excluding puzzles with non-unique solutions or unparseable clue formats, with ground truths derived by the authors' own custom CSP solver. This filtering can bias the result toward puzzles that are easier to parse and have a single solution, and the custom solver is not independently validated. The abstract's phrasing 'On a 35-puzzle retained subset of ZebraLogicBench' is honest, but the significance of the 97% figure should be qualified. The authors should report the full selection criteria, the number of excluded puzzles per size class, validate the solver against the original ZebraLogicBench procedure, and ideally release the solver and the 35 selected puzzle identifiers so the community can reproduce the subset.
- [§4.4, Appendix B] The cross-model claim is unsupported because the Sonnet baseline figures are explicitly flagged as reflecting 'a now-corrected parser limitation (calibrated for Haiku's concise output; misses Sonnet's verbose formatting)' and stated to 'require re-evaluation with a corrected output parser.' Presenting Sonnet Direct/CoT/Self-Refine as 0% with a footnote that the numbers are artifacts of the parser is not a valid comparison. The statement in §4.4 that 'the architecture is model-agnostic' cannot be concluded from these data. The authors should either re-run the Sonnet baselines with a corrected parser or remove the cross-model comparison and the model-agnostic claim.
- [§5.1, Table 9] The component ablation on LGP-6 (6 puzzles, N=3) is the main evidence for the paper's second contribution, but it is very small and all configurations report zero contradictions. This is consistent with the consistency-not-entailment issue: the propagator never fires because the LLM happens to make no internally inconsistent guesses on these six puzzles. The claim that 'MRV guidance is the primary driver' is therefore really a claim about LGP-6 only; on harder or different distributions the safety-net role of contradiction detection could change. The paper acknowledges this in §6, but the contribution statement should be correspondingly hedged.
minor comments (5)
- [§4] The text says '84 problem instances' but the Limitations section says '119 instances'. The difference is the 35 AQUA-RAT problems, but this should be stated explicitly in §4 when the total is introduced.
- [Table 16] The column label 'Verifier type' for SymStep/SymStep+G says 'Propagator'. Since the propagator checks consistency, not entailment, this label should be 'Consistency checker' to avoid aligning with the overstated 'verification' terminology.
- [§3.2] The sentence 'The propagator guarantees consistency with prior accepted deductions, not logical entailment from the clues; a fabricated but locally consistent claim can still be accepted' appears twice in the same paragraph. One occurrence should be removed.
- [§4.8] On AR-LSAT, Direct already achieves 100%, so the comparison against CoT (87%) does not demonstrate that SymStep is necessary; it only shows that CoT is worse. This is acknowledged implicitly but should be stated in the text to avoid overinterpretation.
- [General] The paper uses 'soundness' in several places (e.g., Appendix A, §6) in a way that conflicts with the Appendix F 'Soundness caveat'. Please standardize the terminology: soundness should refer to clue-entailment, not fault-free bookkeeping.
Circularity Check
No significant circularity; the empirical claims are benchmarked externally and the MRV signal is computed from live propagator state, not from ground-truth answers.
full rationale
SymStep's central claims are empirical, and the evaluation pathway does not reduce to the method's own definitions. The MRV hint is computed from the live propagator state after each accepted deduction (Eq. 1), not from ground-truth answer keys, so the 97%/100% results are not predictions of fitted values. The formal results in Appendix F are explicitly internal-consistency invariants, and the paper discloses the load-bearing caveat that consistent-but-incorrect deductions can be accepted (Sec. 3.2 and Appendix F); that is a limitation on the strength of "verification," not a circular derivation. External benchmarks (ZebraLogicBench, AR-LSAT) provide independent test material, and the author-constructed LGP-14/20 benchmarks introduce a mild selection consideration rather than a circular step. The only in-scope concerns—the under-specified uniqueness check in Sec. 2.1 and the fact that the guidance-only ablation still runs the propagator silently—are validity/confound issues, not reductions of the claimed results to their inputs. No load-bearing self-citation chain is present, and no fitted parameter is relabeled as a prediction. Score 1 reflects minor author-constructed benchmark presence; there is no significant circularity.
Assumptions & free parameters
free parameters (1)
- interaction budget T =
not specified
assumptions (3)
- domain assumption Bijective attribute assignment: every attribute value is assigned to exactly one entity, and the propagator enforces this via arc consistency.
- domain assumption LLM DEDUCE assertions are clue-grounded (entailed by the puzzle text), not merely consistent with previously accepted deductions.
- domain assumption The LLM reliably emits exactly one regex-parseable atomic assertion per turn in the DEDUCE format.
Cite this review
Pith. "Pith review of SymStep: Symbolic Step Verification for Logical Reasoning." pith.science (2026). https://pith.science/paper/5FQJ6PUM
@misc{pith2026260723055,
author = {Pith},
title = {Pith review of: SymStep: Symbolic Step Verification for Logical Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/5FQJ6PUM}},
note = {Machine review of arXiv:2607.23055}
}
read the original abstract
Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps. We introduce SymStep: an LLM makes one atomic claim at a time (DEDUCE: Alice, pet, Cat), then a lightweight constraint propagator checks the claim for consistency with prior accepted deductions, rejects contradictions, and cascades implied facts automatically. SymStep+G additionally provides MRV guidance after each accepted step, directing the LLM toward the most constrained unresolved variable. On a 35-puzzle retained subset of ZebraLogicBench, a benchmark of 1,000 Einstein-style logic puzzles, Direct and CoT both achieve 0%, while SymStep+G reaches 97%. On AR-LSAT analytical reasoning problems, SymStep achieves 100% vs. CoT's 87%. On LGP-14, SymStep+G achieves 100% vs. 0% for CoT and Logic-LM, the strongest prior symbolic+LLM baseline we compare against. Ablation studies reveal that MRV guidance is a key mechanism for reducing directionless cycling, while consistency checking provides a safety net against explicit contradictions. Across six benchmarks spanning five task domains, SymStep variants match or exceed every baseline on constraint-dense and arithmetic tasks. Experiments on AQUA-RAT algebra confirm the advantage is constraint-density-specific.
Figures
Reference graph
Works this paper leans on
-
[1]
GPT-4 technical report.arXiv preprint arXiv:2303.08774,
[Achiamet al., 2023 ] Josh Achiam, Steven Adler, Sand- hini Agarwal, Lama Ahmad, Ilge Akkaya, , et al. GPT-4 technical report.arXiv preprint arXiv:2303.08774,
arXiv 2023
-
[5]
[Guoet al., 2025 ] Daya Guo, Dejian Yang, Haowei Zhang, et al. DeepSeek-R1: Incentivizing reasoning capabil- ity in LLMs via reinforcement learning.arXiv preprint arXiv:2501.12948,
arXiv 2025
-
[7]
[Haoet al., 2023 ] Shibo Hao, Yi Gu, Haodi Ma, Joshua Ji- ahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu
Association for Compu- tational Linguistics. [Haoet al., 2023 ] Shibo Hao, Yi Gu, Haodi Ma, Joshua Ji- ahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. Reasoning with language model is planning with world model. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 8154–8173, Singapore, December
2023
-
[8]
[Hoffmann and Nebel, 2001] J¨org Hoffmann and Bernhard Nebel
Associa- tion for Computational Linguistics. [Hoffmann and Nebel, 2001] J¨org Hoffmann and Bernhard Nebel. The FF planning system: Fast plan generation through heuristic search.Journal of Artificial Intelli- gence Research, 14:253–302,
2001
-
[10]
Large language models are zero-shot reasoners
[Kojimaet al., 2022 ] Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. InAd- vances in Neural Information Processing Systems, vol- ume 35, pages 22199–22213
2022
-
[16]
Reflexion: Language agents with verbal reinforcement learning
[Shinnet al., 2023 ] Noah Shinn, Federico Cassano, Ed- ward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. InAdvances in Neu- ral Information Processing Systems, volume 36, pages 8634–8652
2023
-
[21]
Griffiths, Yuan Cao, and Karthik Narasimhan
[Yaoet al., 2023 ] Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate prob- lem solving with large language models. InAdvances in Neural Information Processing Systems 36 (NeurIPS 2023),
2023
-
[22]
Least-to-most prompting enables complex reasoning in large language models
[Zhouet al., 2023 ] Denny Zhou, Nathanael Sch ¨arli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed Chi. Least-to-most prompting enables complex reasoning in large language models. InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5,
2023
Show all 94 references
-
[23]
Natural Program
A Related Work Chain-of-thought and multi-step reasoning.The most popular approach for multi-step reasoning with LLMs is chain-of-thought (CoT) prompting [Weiet al., 2022 ], with extensions ranging from zero-shot variants [Kojimaet al., 2022] and Least-to-Most decomposition [Z...
2022
-
[24]
Bob lives in the blue house
Easy puzzles. E1-Pets: Alice, Bob, Carol;{Red, Blue, Green} × {Cat, Dog, Fish}; 5 clues including “Bob lives in the blue house” and “The person in the red house has a cat.”E2-Jobs: Adam, Beth, Chris;{Tea, Coffee, Water} × {Doctor, Teacher, Engineer}; 5 clues including “The doc...
2023 arXiv
-
[25]
McGuinness, Daniele Nardi, and Peter F
[Baaderet al., 2003 ] Franz Baader, Diego Calvanese, Deb- orah L. McGuinness, Daniele Nardi, and Peter F. Patel- Schneider.The Description Logic Handbook: Theory, Implementation, and Applications. Cambridge Univer- sity Press,
2003
-
[28]
Logical reasoning for task oriented dia- logue systems
[Beygiet al., 2022 ] Sajjad Beygi, Maryam Fazel-Zarandi, Alessandra Cervone, Prakash Krishnan, and Siddhartha Jonnalagadda. Logical reasoning for task oriented dia- logue systems. InProceedings of the Fifth Workshop on e-Commerce and NLP (ECNLP 5), pages 68–79. Asso- ciation f...
2022
-
[29]
Think you have solved direct-answer question answering? try arc-da, the direct-answer AI2 reasoning challenge
[Bhakthavatsalamet al., 2021 ] Sumithra Bhakthavat- salam, Daniel Khashabi, Tushar Khot, Bhavana Dalvi Mishra, Kyle Richardson, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord, and Peter Clark. Think you have solved direct-answer question answering? try arc-da, the direct-...
2021 arXiv
-
[30]
[Blum and Furst, 1995] Avrim Blum and Merrick L. Furst. Fast planning through planning graph analysis. In Proceedings of the 14th International Joint Conference on Artificial Intelligence - Volume 2, IJCAI’95, page 1636–1642, San Francisco, CA, USA,
1995
-
[33]
[Chenet al., 2023 ] Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks.Transactions on Machine Learning Research,
2023
-
[34]
Training verifiers to solve math word problems.ArXiv, abs/2110.14168,
[Cobbeet al., 2021 ] Karl Cobbe, Vineet Kosaraju, Mo Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Rei- ichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems.ArXiv, abs/2110.14168,
2021 arXiv
-
[35]
d’Avila Garcez, Lu´ıs C
[d’Avila Garcezet al., 2009 ] Artur S. d’Avila Garcez, Lu´ıs C. Lamb, and Dov M. Gabbay.Neural-Symbolic Cognitive Reasoning. Cognitive Technologies. Springer,
2009
-
[38]
Pal: program-aided lan- guage models
[Gaoet al., 2023 ] Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. Pal: program-aided lan- guage models. InProceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org,
2023
-
[39]
[Garcez and Lamb, 2023] Artur d’Avila Garcez and Lu´ıs C. Lamb. Neurosymbolic ai: the 3rd wave.Artif. Intell. Rev., 56(11):12387–12406, March
2023
-
[40]
The stable model semantics for logic programming
[Gelfond and Lifschitz, 1988] Michael Gelfond and Vladimir Lifschitz. The stable model semantics for logic programming. InProceedings of the International Logic Programming Conference and Symposium, pages 1070–1080. MIT Press,
1988
-
[42]
[Guoet al., 2025 ] Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu,...
2025
-
[43]
FOLIO: Natural language reasoning with first-order logic
[Hanet al., 2024 ] Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szab ´o, Ekate- rina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong...
2024
-
[44]
[Haoet al., 2023 ] Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu
Association for Computational Linguistics. [Haoet al., 2023 ] Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu. Rea- soning with language model is planning with world model. InProceedings of the 2023 Conference on Em- pirical Methods in Natural La...
2023
-
[45]
[Hendryckset al., 2021 ] Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt
Association for Computational Linguistics. [Hendryckset al., 2021 ] Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring math- ematical problem solving with the MATH dataset. In Advances in Neural Inform...
2021
-
[46]
The ff planning system: fast plan generation through heuristic search
[Hoffmann and Nebel, 2001] J¨org Hoffmann and Bernhard Nebel. The ff planning system: fast plan generation through heuristic search. 14(1):253–302, May
2001
-
[47]
Thought cloning: Learning to think while acting by imitating human thinking
[Hu and Clune, 2023] Shengran Hu and Jeff Clune. Thought cloning: Learning to think while acting by imitating human thinking. InAdvances in Neural Information Processing Systems, volume 36, pages 44451–44469,
2023
-
[48]
Inner monologue: Embodied reasoning through plan- ning with language models
[Huanget al., 2023 ] Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Tomas Jackson, Noah Brown, Linda Luu, Sergey Levine, Karol Hausman, and brian ichter. Inner monologue: ...
2023
-
[49]
Understanding the planning of LLM agents: A survey.arXiv preprint arXiv:2402.02716,
[Huanget al., 2024 ] Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, Defu Lian, Yasheng Wang, Ruiming Tang, and Enhong Chen. Understanding the planning of LLM agents: A survey.arXiv preprint arXiv:2402.02716,
2024 arXiv
-
[50]
Joshi, Kyle Jeffrey, Rosario Jauregui Ruano, Jasmine Hsu, Keerthana Gopalakrishnan, Byron David, Andy Zeng, and Chuyuan Kelly Fu
[Ichteret al., 2023 ] Brian Ichter, Anthony Brohan, Yev- gen Chebotar, Chelsea Finn, Karol Hausman, Alexan- der Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, Dmitry Kalashnikov, Sergey Levine, Yao Lu, Carolina Parada, Kanishka Rao, Pierre Ser- manet, Ale...
2023
-
[51]
Position: Llms can’t plan, but can help plan- ning in llm-modulo frameworks
[Kambhampatiet al., 2024 ] Subbarao Kambhampati, Karthik Valmeekam, Lin Guan, Mudit Verma, Kaya Stechly, Siddhant Bhambri, Lucas Saldyt, and Anil Murthy. Position: Llms can’t plan, but can help plan- ning in llm-modulo frameworks. InProceedings of the 41st International Confer...
2024
-
[52]
Large language models are zero-shot reasoners
[Kojimaet al., 2022 ] Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. InAd- vances in Neural Information Processing Systems, vol- ume 35, pages 22199–22213,
2022
-
[53]
[Leet al., 2022 ] Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven C.H. Hoi. CodeRL: Mastering code generation through pretrained models and deep reinforcement learning. InAdvances in Neural Information Processing Systems 35 (NeurIPS 2022),
2022
-
[54]
Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals
[Liet al., 2022 ] Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, R ´emi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Mas- son d’Autume, Igor Babuschkin, Xinyun Chen, Po- Sen Huang, Johannes...
2022
-
[55]
Code as policies: Language model pro- grams for embodied control
[Lianget al., 2023 ] Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model pro- grams for embodied control. In2023 IEEE International Conference on Robotics and Automation (ICRA),
2023
-
[56]
Let’s Verify Step by Step
[Lightmanet al., 2024 ] Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s Verify Step by Step. In International Conference on Learning Representations,
2024
-
[57]
Zebralogic: On the scaling limits of llms for logical reasoning
[Linet al., 2025 ] Bill Yuchen Lin, Ronan Le Bras, Kyle Richardson, Ashish Sabharwal, Radha Poovendran, Pe- ter Clark, and Yejin Choi. Zebralogic: On the scaling limits of llms for logical reasoning. InForty-second International Conference on Machine Learning, ICML 2025, Vanco...
2025
-
[58]
Program induction by rationale generation: Learning to solve and explain algebraic word problems
[Linget al., 2017 ] Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. Program induction by rationale generation: Learning to solve and explain algebraic word problems. InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL 2017), p...
2017
-
[60]
[Liuet al., 2023 ] Bo Liu, Yuqian Jiang, Xiaohan Zhang, Qiang Liu, Shiqi Zhang, Joydeep Biswas, and Pe- ter Stone
arXiv:2306.03872. [Liuet al., 2023 ] Bo Liu, Yuqian Jiang, Xiaohan Zhang, Qiang Liu, Shiqi Zhang, Joydeep Biswas, and Pe- ter Stone. Llm+p: Empowering large language mod- els with optimal planning proficiency.arXiv preprint arXiv:2304.11477,
2023 arXiv
-
[61]
Agentbench: Eval- uating llms as agents
[Liuet al., 2024 ] Xiao Liu, Hao Yu, Hanchen Zhang, Yi- fan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xi- ang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, an...
2024
-
[62]
Faithful chain-of-thought rea- soning
[Lyuet al., 2023 ] Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch. Faithful chain-of-thought rea- soning. InProceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd C...
2023
-
[63]
[Mackworth, 1977] Alan K
Association for Computational Linguistics. [Mackworth, 1977] Alan K. Mackworth. Consistency in networks of relations.Artificial Intelligence, 8(1):99– 118,
1977
-
[64]
Self-refine: Iterative refinement with self-feedback
[Madaanet al., 2023 ] Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark....
2023
-
[65]
McCarthy and P
[McCarthy and Hayes, 1987] J. McCarthy and P. J. Hayes. Some philosophical problems from the standpoint of arti- ficial intelligence, page 26–45. Morgan Kaufmann Pub- lishers Inc., San Francisco, CA, USA,
1987
-
[69]
Association for Computational Linguis- tics. [OpenAIet al., 2024 ] OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Al- tenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji,...
2024
-
[70]
Logic-LM: Empow- ering large language models with symbolic solvers for faithful logical reasoning
[Panet al., 2023 ] Liangming Pan, Alon Albalak, Xinyi Wang, and William Yang Wang. Logic-LM: Empow- ering large language models with symbolic solvers for faithful logical reasoning. InFindings of the Associa- tion for Computational Linguistics: EMNLP 2023, pages 3806–3824,
2023
-
[71]
Multi-logieval: Towards evaluating multi-step logical reasoning ability of large language models
[Patelet al., 2024 ] Nisarg Patel, Mohith Kulkarni, Mihir Parmar, Aashna Budhiraja, Mutsumi Nakamura, Neeraj Varshney, and Chitta Baral. Multi-logieval: Towards evaluating multi-step logical reasoning ability of large language models. InProceedings of the 2024 Confer- ence on ...
2024
-
[72]
Mea- suring and narrowing the compositionality gap in lan- guage models
[Presset al., 2023 ] Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah Smith, and Mike Lewis. Mea- suring and narrowing the compositionality gap in lan- guage models. InFindings of the Association for Com- putational Linguistics: EMNLP 2023, pages 5687–5711, Singapore, December
2023
-
[73]
[Puiget al., 2018 ] Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Tor- ralba
Association for Computa- tional Linguistics. [Puiget al., 2018 ] Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Tor- ralba. VirtualHome: Simulating household activities via programs. InProceedings of the IEEE Conference on Computer Vision...
2018
-
[77]
Textgraphs 2024 shared task on text-graph representations for knowledge graph question answering
[Sakhovskiyet al., 2024 ] Andrey Sakhovskiy, Mikhail Salnikov, Irina Nikishina, Aida Usmanova, Angelie Kraft, Cedric M ¨oller, Debayan Banerjee, Junbo Huang, Longquan Jiang, Rana Abdullah, Xi Yan, Dmitry Ustalov, Elena Tutubalina, Ricardo Usbeck, and Alexander Panchenko. Textg...
2024
-
[78]
[Schicket al., 2023 ] Timo Schick, Jane Dwivedi-Yu, Roberto Dess ´ı, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom
Association for Computational Linguistics. [Schicket al., 2023 ] Timo Schick, Jane Dwivedi-Yu, Roberto Dess ´ı, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: language models can teach themselves to use tools. In...
2023
-
[79]
Hug- ginggpt: Solving ai tasks with chatgpt and its friends in hugging face
[Shenet al., 2023 ] Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. Hug- ginggpt: Solving ai tasks with chatgpt and its friends in hugging face. InAdvances in Neural Information Pro- cessing Systems, volume 36, pages 38154–38180,
2023
-
[80]
Reflexion: language agents with verbal reinforcement learning
[Shinnet al., 2023 ] Noah Shinn, Federico Cassano, Ash- win Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: language agents with verbal reinforcement learning. InAdvances in Neural Information Processing Systems, volume 36, pages 8634–8652,
2023
-
[81]
ALFRED: A benchmark for interpreting grounded instructions for everyday tasks
[Shridharet al., 2020 ] Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. ALFRED: A benchmark for interpreting grounded instructions for everyday tasks. In2020 IEEE/CVF Conference on Com- puter Vision a...
2020
-
[83]
Tenenbaum, Leslie Pack Kaelbling, and Michael Katz
[Silveret al., 2024 ] Tom Silver, Soham Dan, Kavitha Srinivas, Joshua B. Tenenbaum, Leslie Pack Kaelbling, and Michael Katz. Generalized planning in PDDL do- mains with pretrained large language models. InPro- ceedings of the AAAI Conference on Artificial Intelli- gence, volum...
2024
-
[84]
ProgPrompt: Generating situated robot task plans us- ing large language models
[Singhet al., 2023 ] Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Trem- blay, Dieter Fox, Jesse Thomason, and Animesh Garg. ProgPrompt: Generating situated robot task plans us- ing large language models. In2023 IEEE International Conference o...
2023
-
[85]
Smolensky
[Smolensky, 1990] P. Smolensky. Tensor product variable binding and the representation of symbolic structures in connectionist systems.Artif. Intell., 46(1–2):159–216, November
1990
-
[87]
Sadler, Wei-Lun Chao, and Yu Su
[Songet al., 2023 ] Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M. Sadler, Wei-Lun Chao, and Yu Su. Llm-planner: Few-shot grounded planning for embod- ied agents with large language models. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV...
2023
-
[88]
Beyond the imitation game: Quantifying and extrapolating the capabilities of lan- guage models.arXiv preprint arXiv:2206.04615,
[Srivastavaet al., 2022 ] Aarohan Srivastava, Abigail Ras- togi, Abhishek Rao, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of lan- guage models.arXiv preprint arXiv:2206.04615,
2022 arXiv
-
[89]
Griffiths
[Sumerset al., 2024 ] Theodore Sumers, Shunyu Yao, Karthik R Narasimhan, and Thomas L. Griffiths. Cog- nitive Architectures for Language Agents.Transactions on Machine Learning Research,
2024
-
[90]
Le, Denny Zhou, William Fedus, et al
[Suzgunet al., 2022 ] Mirac Suzgun, Nathan Scales, Nathanael Sch ¨arli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V . Le, Denny Zhou, William Fedus, et al. Challenging BIG-Bench tasks and whether chain-of-thought can solve them.arXiv preprint arXiv...
2022 arXiv
-
[91]
ProofWriter: Generating implications, proofs, and abductive statements over natural language
[Tafjordet al., 2021 ] Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. ProofWriter: Generating implications, proofs, and abductive statements over natural language. InFindings of the Association for Computational Lin- guistics: ACL-IJCNLP 2021, pages 3621–3634, Online, August
2021
-
[92]
[Talmoret al., 2019 ] Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant
Association for Computational Linguis- tics. [Talmoret al., 2019 ] Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Common- senseQA: A question answering challenge targeting commonsense knowledge. InProceedings of the 2019 Conference of the North American Ch...
2019
-
[94]
LLMs still can’t plan; can LRMs? a preliminary evaluation of OpenAI’s o1 on PlanBench.arXiv preprint arXiv:2409.13373,
[Valmeekamet al., 2024 ] Karthik Valmeekam, Kaya Stechly, and Subbarao Kambhampati. LLMs still can’t plan; can LRMs? a preliminary evaluation of OpenAI’s o1 on PlanBench.arXiv preprint arXiv:2409.13373,
2024 arXiv
-
[95]
Towards data- and knowledge-driven artificial intelli- gence: A survey on neuro-symbolic computing.IEEE Transactions on Pattern Analysis and Machine Intelli- gence,
[Wanget al., 2024 ] Wenguan Wang, Yi Yang, and Fei Wu. Towards data- and knowledge-driven artificial intelli- gence: A survey on neuro-symbolic computing.IEEE Transactions on Pattern Analysis and Machine Intelli- gence,
2024
-
[96]
[Weiet al., 2022 ] Jason Wei, Xuezhi Wang, Dale Schuur- mans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V
arXiv:2210.15889. [Weiet al., 2022 ] Jason Wei, Xuezhi Wang, Dale Schuur- mans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V . Le, and Denny Zhou. Chain-of-thought prompt- ing elicits reasoning in large language models. InAd- vances in Neural Information Processing Sys...
2022 arXiv
-
[97]
Neuro-symbolic relation extraction
[Yanet al., 2025 ] Xi Yan, Aida Usmanova, Cedric M¨oller, Patrick Westphal, and Ricardo Usbeck. Neuro-symbolic relation extraction. InHandbook on Neurosymbolic AI and Knowledge Graphs, volume 400 ofFrontiers in Arti- ficial Intelligence and Applications, pages 550–576. IOS Press,
2025
-
[98]
[Yanget al., 2018 ] Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Meth- ods in ...
2018
-
[99]
InterCode: Standardiz- ing and benchmarking interactive coding with execution feedback
[Yanget al., 2023 ] John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao. InterCode: Standardiz- ing and benchmarking interactive coding with execution feedback. InAdvances in Neural Information Processing Systems 36 (NeurIPS 2023),
2023
-
[100]
Griffiths, Yuan Cao, and Karthik Narasimhan
[Yaoet al., 2023a ] Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate prob- lem solving with large language models. InAdvances in Neural Information Processing Systems 36 (NeurIPS 2023),
2023
-
[101]
HellaSwag: Can a machine really finish your sentence?arXiv preprint arXiv:1905.07830,
[Zellerset al., 2019 ] Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence?arXiv preprint arXiv:1905.07830,
2019 arXiv
-
[102]
Large language models as commonsense knowl- edge for large-scale task planning
[Zhaoet al., 2023 ] Zirui Zhao, Wee Sun Lee, and David Hsu. Large language models as commonsense knowl- edge for large-scale task planning. InAdvances in Neu- ral Information Processing Systems, volume 36, pages 31967–31987,
2023
-
[103]
Le, and Ed H
[Zhouet al., 2023 ] Denny Zhou, Nathanael Sch ¨arli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V . Le, and Ed H. Chi. Least-to-most prompting enables complex reasoning in large language models. InThe Eleventh Internation...
2023
-
[1965]
Wino- Grande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106,
[Sakaguchiet al., 2021 ] Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Wino- Grande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106,
2021
-
[1971]
[Fuggitti and Chakraborti, 2023] Francesco Fuggitti and Tathagata Chakraborti
Morgan Kaufmann Publishers Inc. [Fuggitti and Chakraborti, 2023] Francesco Fuggitti and Tathagata Chakraborti. NL2LTL – a python package for converting natural language (NL) instructions to linear temporal logic (LTL) formulas. InProceedings of the AAAI Conference on Artificia...
2023
-
[1977]
Self-refine: Iterative refinement with self-feedback
[Madaanet al., 2023 ] Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prab- humoye, Yiming Yang, et al. Self-refine: Iterative refinement with self-feedback. InAdvances in Neural Information Processing System...
2023
-
[1987]
PDDL – the planning domain definition language
[McDermottet al., 1998 ] Drew McDermott, Malik Ghal- lab, Adele Howe, Craig Knoblock, Ashwin Ram, Manuela Veloso, Daniel Weld, and David Wilkins. PDDL – the planning domain definition language. Tech- nical Report CVC TR-98-003 / DCS TR-1165, Yale Cen- ter for Computational Vis...
1998
-
[1988]
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies.Transactions of the Association for Computational Linguistics, 9:346–361,
[Gevaet al., 2021 ] Mor Geva, Daniel Khashabi, Elad Se- gal, Tushar Khot, Dan Roth, and Jonathan Berant. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies.Transactions of the Association for Computational Linguistics, 9:346–361,
2021
-
[1990]
Scaling llm test-time com- pute optimally can be more effective than scaling model parameters.ArXiv, abs/2408.03314,
[Snellet al., 2024 ] Charlie Victor Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time com- pute optimally can be more effective than scaling model parameters.ArXiv, abs/2408.03314,
2024 arXiv
-
[1995]
[Bollackeret al., 2008 ] Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor
Morgan Kaufmann Publishers Inc. [Bollackeret al., 2008 ] Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. Freebase: a collaboratively created graph database for structuring human knowledge. InProceedings of the 2008 ACM SIGMOD International Conferen...
2008
-
[1998]
On the compilability and expressive power of propositional planning formalisms
[Nebel, 2000] Bernhard Nebel. On the compilability and expressive power of propositional planning formalisms. 12(1):271–315, May
2000
-
[2000]
LINC: A neurosymbolic approach for logical reasoning by combining language models with first-order logic provers
[Olaussonet al., 2023 ] Theo Olausson, Alex Gu, Ben Lip- kin, Cedegao Zhang, Armando Solar-Lezama, Joshua Tenenbaum, and Roger Levy. LINC: A neurosymbolic approach for logical reasoning by combining language models with first-order logic provers. InProceedings of the 2023 Conf...
2023
-
[2001]
Position: LLMs can’t plan, but can help planning in LLM-Modulo frameworks.arXiv preprint arXiv:2402.01817,
[Kambhampatiet al., 2024 ] Subbarao Kambhampati, Karthik Valmeekam, Lin Guan, Mudit Verma, Kaya Stechly, Siddhant Bhambri, Lucas Saldyt, and Anil Murthy. Position: LLMs can’t plan, but can help planning in LLM-Modulo frameworks.arXiv preprint arXiv:2402.01817,
2024 arXiv
-
[2003]
Semantic parsing on freebase from question-answer pairs
[Berantet al., 2013 ] Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. Semantic parsing on freebase from question-answer pairs. InProceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1533–1544,
2013
-
[2008]
[Caiet al., 2024 ] Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou
Association for Computing Machinery. [Caiet al., 2024 ] Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou. Large language models as tool makers. InInternational Conference on Learn- ing Representations, volume 2024, pages 54067–54089,
2024
-
[2009]
Fikes and Nils J
[Fikes and Nilsson, 1971] Richard E. Fikes and Nils J. Nilsson. Strips: a new approach to the application of theorem proving to problem solving. InProceedings of the 2nd International Joint Conference on Artificial In- telligence, IJCAI’71, page 608–620, San Francisco, CA, USA,
1971
-
[2010]
[Robinson, 1965] J. A. Robinson. A machine-oriented logic based on the resolution principle.J. ACM, 12(1):23–41, January
1965
-
[2013]
Graph of thoughts: Solving elaborate problems with large language models
[Bestaet al., 2024 ] Maciej Besta, Nils Blach, Ales Ku- bicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. Graph of thoughts: Solving elaborate problems with large language mod...
2024
-
[2017]
Deductive verification of chain-of-thought reason- ing
[Linget al., 2023 ] Zhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang, Mingu Lee, Roland Memisevic, and Hao Su. Deductive verification of chain-of-thought reason- ing. InAdvances in Neural Information Processing Sys- tems 36 (NeurIPS 2023),
2023
-
[2018]
The lama planner: guiding cost-based any- time planning with landmarks
[Richter and Westphal, 2010] Silvia Richter and Matthias Westphal. The lama planner: guiding cost-based any- time planning with landmarks. 39(1):127–177, Septem- ber
2010
-
[2019]
LLaMA: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971,
[Touvronet al., 2023 ] Hugo Touvron, Thibaut Lavril, Gau- tier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, Armand Joulin, Edouard Grave, and Guillaume Lample. LLaMA: Open and efficient foundation ...
2023 arXiv
-
[2020]
ALFWorld: Aligning text and embodied environments for interactive learning
[Shridharet al., 2021 ] Mohit Shridhar, Xingdi Yuan, Marc-Alexandre C ˆot´e, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. ALFWorld: Aligning text and embodied environments for interactive learning. In International Conference on Learning Representations (ICLR 2021),
2021
-
[2021]
d’Avila Garcez and Lu´ıs C
[d’Avila Garcez and Lamb, 2023] Artur S. d’Avila Garcez and Lu´ıs C. Lamb. Neurosymbolic AI: The 3rd wave. Artificial Intelligence Review, 56:12387–12406,
2023
-
[2022]
Let’s verify step by step
[Lightmanet al., 2024 ] Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. InIn- ternational Conference on Learning Representations,
2024
-
[2025]
FOLIO: Natural lan- guage reasoning with first-order logic
[Hanet al., 2024 ] Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, Ar- man Cohan, Dragomir Radev, et al. FOLIO: Natural lan- guage reasoning with first-order logic. InProceedings of the 2024 Conference on Empirical Methods in Natu- ral Lang...
2024
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.