Pith. sign in

REVIEW 3 major objections 6 minor 78 references

Distilling Answer Set Programming Theories from Large Language Models

T0 review · 3 major / 6 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read Frontier LLM agents can author complete ASP reasoning theories from an empty file in one hour, matching or beating handwritten baselines on VQA benchmarks.

desk verdict Solid empirical result: frontier agents can author full ASP reasoning theories from an empty file under a fixed solver-in-the-loop protocol; the main bound is oracle upstream facts, not a broken experiment. read the letter →

arxiv 2607.28086 v1 pith:6S3TMU5T submitted 2026-07-30 cs.AI

classification cs.AI
keywords AnswerSetProgrammingLargeLanguageModelsNeurosymbolicReasoningLLMAgentsVisualQuestionAnsweringTheoryDistillationclingo
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Writing Answer Set Programming theories by hand is slow and does not transfer across domains. This paper asks whether a large language model, given only an empty file, a fixed minimal prompt, and shell access to a solver, can iteratively author a full reasoning theory within a one-hour budget. On three Visual Question Answering benchmarks—CLEVR, GQA, and CLEVRER—the agent receives ground-truth scene and question facts and must write the rules that turn those facts into answers. Three of four frontier models reach 100% on CLEVR, 92.8–98.8% on GQA (above the handwritten reference), and 92.7–95.3% on CLEVRER. A fourth frontier model collapses on the harder real-image setting, and smaller models mostly fail to produce usable theories. The result matters because it treats the symbolic program itself as the thing the model must build, not a one-shot answer, under a protocol that is the same across datasets.

What carries the argument

The distillation protocol: a sandboxed agent loop in which the model may only read training examples, edit a single theory file, and call clingo or a linter, starting from an empty file and a fixed minimal prompt, until it stops or a one-hour cap fires; quality is measured by held-out validation accuracy of the final theory.

What would settle it

Run the same empty-file, one-hour protocol on GQA or CLEVRER but replace ground-truth scene and question facts with outputs from a real vision module and learned parser; if frontier-model theories fall far below the reported 93%+ band, the central claim does not transfer beyond oracle upstream inputs.

Watch

Extended reading notes

Core claim

Under a fixed, dataset-agnostic agent harness with the solver in the loop, three of four frontier models distill complete ASP theories from scratch that reach ceiling accuracy on CLEVR, meet or exceed the handwritten GQA reference, and score above 92% on CLEVRER, while sub-27B models and one frontier model largely fail on coverage or tool use.

Load-bearing premise

The agent is handed perfect scene and question facts from each dataset’s own annotations, so it only has to write the reasoning rules—not deal with noisy vision or imperfect parsers.

Editorial extensions

If this is right

  • Hand-authoring ASP theories for VQA-style reasoning can be replaced, at frontier scale, by a fixed agent loop with a solver in the loop.
  • Reference theories from other domains are optional and can hurt some models; the baseline empty-file setting is already enough for three frontier models.
  • There is a sharp capability threshold: below roughly 27B parameters, models fail on tool format, ASP syntax, or stall before authoring a theory.
  • Distilled theories can be plugged in as the symbolic module of existing neurosymbolic VQA pipelines and used to supervise vision or parser training from solver feedback.
  • The agent loop is not uniformly helpful: for at least one frontier model on GQA, a one-shot theory outperforms iterative editing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If upstream perception is the real bottleneck, the practical next product is a closed loop that trains vision against a frozen distilled theory rather than further theory authoring.
  • The GPT-5 GQA collapse and reference-induced shortening suggest context budget and coverage style, not raw scale alone, decide whether the loop helps or hurts.
  • A similar empty-file solver-in-the-loop protocol could test whether agents can distill full theories for non-VQA ASP domains (planning, configuration, diagnosis) without per-task templates.
  • Releasing the theories enables direct comparison of human vs model rule style—length is not coverage—and audits of which operators remain systematically missing.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper studies whether LLM agents can author complete Answer Set Programming (ASP) reasoning theories from scratch under a fixed, dataset-agnostic protocol: empty theory.lp, a minimal prompt with no ASP primer, shell access to clingo (solve/lint only), and a 1-hour autonomous budget in a Docker sandbox. The application is VQA reasoning on CLEVR, GQA, and CLEVRER, with ground-truth scene and question facts supplied from each dataset’s annotations; the agent authors only the interpreter theory T. Nine models are evaluated (four frontier, two mid-tier, three smaller open-weights), with N=3 seeds per configuration, held-out val/200 scoring under strict/brave entailment, a reference-theory ablation (B∈{0,1,2}), a one-shot non-agent baseline, tool-use profiles, and a failure-mode taxonomy. Three frontier models reach 100% on CLEVR, 92.8–98.8% on GQA (above the 77.5% handwritten reference), and 92.7–95.3% on CLEVRER at B=0; GPT-5 is strong on CLEVR but collapses on GQA (41.8%), and references help the other three little while hurting GPT-5. Code, prompts, and distilled theories are released.

Significance. If the results hold, the paper provides concrete evidence that frontier LLM agents can synthesize non-trivial, inspectable ASP theories end-to-end under a solver-in-the-loop harness, without per-operator templates or human rule scaffolding. That is a useful empirical contribution to neurosymbolic reasoning and to LLM-agent evaluation: the unit of work is a full theory across many edit cycles, not single-shot program fragments. Strengths that should be credited include the fixed harness/prompt across all configurations, held-out validation never visible to the agent, independent train/val splits per seed, the one-shot baseline that isolates when the agent loop helps versus hurts (notably GPT-5 on GQA), the reference ablation, the failure taxonomy (parse/no-answer/semantic and session-level modes), accuracy-growth curves, and full release of code, prompts, and theories. The main bound on significance is scope: perception and parsing are oracle, so the result is interpreter synthesis for a known DSL under clean I/O rather than full end-to-end VQA robustness.

major comments (3)
  1. [§3.2, §5.2, Table 1] §3.2 and §5.2 / Table 1: The GQA handwritten reference reaches only 77.5% with 57 rules, while distilled theories reach 92.8–98.8% with 65–343 rules. The manuscript repeatedly frames this as meeting or exceeding the “handwritten GQA ceiling.” That wording overstates the comparison unless the paper shows the handwritten theory is near-complete for the operator/attribute vocabulary. As written, the gap is more naturally read as incomplete human coverage of GQA’s large schema than as surpassing a strong human theory. Please quantify operator/question-shape coverage of the handwritten GQA theory versus the distilled ones, and rephrase claims of “exceeding the handwritten ceiling” accordingly (e.g., completing coverage left incomplete by the reference).
  2. [Abstract, §1, RQ1, §3.1–3.2] Abstract, §1, and RQ1 vs. §3.1–3.2 / Fig. 2: The central experimental claim is well supported for authoring T given oracle s_asp and q_asp. However, phrases such as “complete and correct theories” and the neurosymbolic-VQA framing can be read as end-to-end competence. The protocol never stresses T under noisy or incomplete upstream facts; Conclusion correctly lists coupling to learned perception/parsers as future work. Tighten abstract/intro/RQ1 language so the load-bearing claim matches the measured setting (interpreter synthesis under clean parser outputs), and state explicitly that reported acc_strict does not transfer by default to imperfect scene/question facts.
  3. [§5.1, Table 1, Appendix C] §5.1 and Table 1: N=3 independent seeds is understandable given cost, and frontier B=0 stds are often small. Several load-bearing secondary claims rest on much noisier cells (GPT-5 GQA 41.8±9.8; GPT-5 CLEVRER B=2 67.3±13.2; Flash and qwen3.6-27b with ±20–50 pp). For those comparisons—especially “agent loop is net-negative for GPT-5 on GQA” versus the one-shot baseline in Appendix C—please either add seeds, report confidence intervals / pairwise tests, or clearly mark which conclusions are qualitative. As is, the GPT-5 regression and sub-frontier threshold claims are directionally plausible but statistically thin.
minor comments (6)
  1. [Table 2, Table 3] Table 3 vs. Table 2: Theory size (rules) is reported as mean over non-empty theories, while failure rates include empty/broken sessions. A short note that size is conditional on producing a theory would avoid over-reading GPT-5’s short GQA theory as efficient coverage.
  2. [Figure 4, Appendix G] Figure 4: Accuracy-growth curves are informative; consider marking the handwritten GQA 77.5% line and noting edit counts at plateau for GPT-5 to support the soft-fail-plateau discussion in Appendix G.
  3. [Appendix A, §5.4, §6] Appendix A model sizes mix exact MoE counts with “community estimates.” Flag estimated sizes more visibly in the main text when discussing the “27B boundary,” since that threshold is used interpretively in §5.4–6.
  4. [§2] Related Work: The distinction from Eiter et al. (2024) (rule-by-rule vs. whole-theory, no per-question-shape template) is clear; a one-sentence comparison of evaluation protocol (held-out acc vs. assembly completeness) would help readers place the contribution.
  5. [Abstract, §5.2] Minor typos/grammar: abstract “we nine different models”; §5.2 “T ool use”; occasional missing spaces after figure/table references. Sweep for these before camera-ready.
  6. [§3.1, Table 1] Scoring details (brave vs. strict singleton for yes/no) are in §3.1 and the system prompt; a brief restatement next to Table 1 would make acc_strict self-contained for readers who skip appendices.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: held-out val accuracy against dataset ground truth is independent of the authored theory and of optional cross-dataset references.

full rationale

The paper’s load-bearing claims are experimental (Table 1, B=0): with an empty theory.lp, a fixed minimal prompt, clingo in the loop, and a 1-hour cap, frontier agents author ASP theories scored by acc_strict on held-out val/200. Validation examples are staged outside the sandbox and never visible to the model (Sec. 4.1–4.3); answers are checked by brave/strict entailment against each dataset’s own ground-truth labels, not quantities defined from T. Optional B≥1 inputs are read-only handwritten theories from other datasets only; the target dataset’s own theory is never supplied (Sec. 5.1, Appendix H). Prior Eiter et al. citations supply those optional references and the handwritten ceilings (CLEVR 100%, GQA 77.5%), which are comparison baselines, not premises that force the distilled accuracies. Growth curves, failure taxonomy, one-shot non-agent baseline (Appendix C), and released artifacts further separate authoring dynamics from the metric. Nothing in the derivation reduces a claimed prediction to a fitted input or to a self-definitional identity. Scope limits (oracle scene_asp/question_asp) affect transfer, not circularity of the stated distillation results.

Assumptions & free parameters 3 free parameters · 5 assumptions · 1 invented entities

This is an empirical systems paper. Load-bearing premises are methodological choices (empty-file start, fixed prompt, 1-hour cap, ground-truth ASP facts, N=3, val/200, brave/strict scoring), not fitted physical constants or invented particles. No free parameters are fit to produce the headline accuracies; model identity and B∈{0,1,2} are experimental factors.

free parameters (3)
  • time budget τ = 1 hour = 1 hour
    Hard cap on agent wall-clock time per sample; directly bounds how much theory editing is possible and is chosen by the experimenters.
  • train/val split sizes (100 train + 200 val) = 100 / 200
    Fixed evaluation design drawn from a 10× pool; accuracy numbers depend on this choice though not fitted to maximize a claim.
  • N = 3 samples per configuration = 3
    Replication count chosen by authors; variance estimates and means rest on this small N.
assumptions (5)
  • standard math Stable-model (answer set) semantics and clingo enumeration correctly implement the intended ASP reasoning for ans/1 scoring.
    Background from Gelfond & Lifschitz 1988 and Gebser et al. 2019; assumed throughout Sections 3–5.
  • domain assumption Dataset-provided scene graphs / functional programs can be faithfully compiled into scene_asp and question_asp facts that fully determine the reasoning problem.
    Stated in Sections 3.1–3.2; perception and parsing are out of scope by design.
  • domain assumption Brave entailment for open-ended answers and singleton commitment for yes/no is the right strict accuracy metric for theory quality.
    Scoring definition in Section 3.1 and system prompt; defines acc_strict reported in Table 1.
  • ad hoc to paper A single fixed minimal prompt with no ASP primer is a fair, dataset-agnostic test of model capability rather than prompt engineering.
    Section 4.3 explicitly holds prompts constant so differences ‘reflect model capability, not prompt engineering.’
  • ad hoc to paper One-hour autonomous agent sessions in a Docker sandbox with restricted bash (solve/lint only) adequately represent ‘distillation from scratch.’
    Protocol in Section 4; excludes human-in-the-loop edits and longer runs.
invented entities (1)
  • Dataset-agnostic ASP theory distillation protocol (empty theory.lp + fixed harness + solver-in-the-loop) independent evidence
    purpose: Operationalize whole-theory authoring as a comparable experimental unit across models and datasets.
    The protocol is the paper’s main methodological object (Figure 3, Algorithm 1). It is a procedure, not a physical entity; independent evidence is the released code and multi-model results.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Distilling Answer Set Programming Theories from Large Language Models." pith.science (2026). https://pith.science/paper/6S3TMU5T

@misc{pith2026260728086,
  author       = {Pith},
  title        = {Pith review of: Distilling Answer Set Programming Theories from Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6S3TMU5T}},
  note         = {Machine review of arXiv:2607.28086}
}
read the original abstract

Writing Answer Set Programming (ASP) theories from scratch is a difficult and time-consuming task. We take a neurosymbolic approach to study whether a model can distill complete and correct theories, given a fixed agent harness with the solver in the loop. The protocol is dataset-agnostic: with a single prompt and an empty file as the starting point the model is given a 1-hour time limit to derive a complete theory. We chose VQA as the application domain, three benchmarks (CLEVR, GQA, CLEVRER), as these are publicly available and non-trivial. In order to study the model scale required for solving this task we nine different models: four frontier (Claude Sonnet 4.6, Claude Opus 4.7, GPT-5, DeepSeek V4 Pro), two mid-tier (DeepSeek V4 Flash, gpt-oss-120b), and three open-weights (qwen3.6-27b, gpt-oss-20b, qwen3.5-9b). Three of four frontier models reach 100% on CLEVR and 92.8%-98.8% on GQA; on CLEVRER, Sonnet, Opus, DeepSeek V4 Pro score 92.7%-95.3%. GPT-5 reaches 98.7% on CLEVR but drops to 41.8% on GQA and to 86.7% on CLEVRER. Adding handwritten reference theories from other datasets moves the other three frontier models by at most +/-3.4 pp but reduces GPT-5's accuracy by 3-19 pp. We release the code, prompts, and theories distilled.

Figures

Figures reproduced from arXiv: 2607.28086 by the authors.

Figure 1
Figure 1. Representative examples of the three datasets. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. One GQA example end-to-end, against the scene in Figure [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. System overview. Three inputs (dataset D, B reference theories from other datasets, model M) enter a sandbox in which M iteratively picks a tool (read, edit, write, glob, grep) to read a training example or reference, edit the theory T, or call the ASP solver on T. The loop runs until M self-stops or the 1-hour cap fires; the output is the final theory T, scored on a held-out validation set outside the sandbox. 4. D… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Accuracy-growth curves: acc strict on val/200 as a function of edits to theory, B = 0 condition, frontier models only. Each curve is the mean across 3 samples; the shaded band is ±1 SEM (standard error of the mean, = sample std/ √ 3). CLEVR (left) is solved by every mo…
Figure 5
Figure 5. Figure 5: Tool use distribution per model, summed across every completed sample. Each [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

78 extracted references · 6 canonical work pages

  1. [1]

    The Stable Model Semantics for Logic Programming , booktitle =

    Michael Gelfond and Vladimir Lifschitz , editor =. The Stable Model Semantics for Logic Programming , booktitle =

  2. [2]

    Theory Pract

    Martin Gebser and Roland Kaminski and Benjamin Kaufmann and Torsten Schaub , title =. Theory Pract. Log. Program. , volume =. 2019 , doi =

  3. [3]

    Lawrence Zitnick and Devi Parikh , title =

    Stanislaw Antol and Aishwarya Agrawal and Jiasen Lu and Margaret Mitchell and Dhruv Batra and C. Lawrence Zitnick and Devi Parikh , title =. 2015. 2015 , doi =

  4. [4]

    Making the

    Yash Goyal and Tejas Khot and Douglas Summers. Making the. 2017. 2017 , doi =

  5. [5]

    Justin Johnson and Bharath Hariharan and Laurens van der Maaten and Li Fei. 2017. 2017 , doi =

  6. [6]

    Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations , journal =

    Ranjay Krishna and Yuke Zhu and Oliver Groth and Justin Johnson and Kenji Hata and Joshua Kravitz and Stephanie Chen and Yannis Kalantidis and Li. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations , journal =. 2017 , doi =

  7. [7]

    Hudson and Christopher D

    Drew A. Hudson and Christopher D. Manning , title =. 2019 , doi =

  8. [8]

    Tenenbaum , title =

    Kexin Yi and Chuang Gan and Yunzhu Li and Pushmeet Kohli and Jiajun Wu and Antonio Torralba and Joshua B. Tenenbaum , title =. 8th International Conference on Learning Representations,. 2020 , url =

Show all 78 references
  1. [9]

    Jacob Andreas and Marcus Rohrbach and Trevor Darrell and Dan Klein , title =. 2016. 2016 , doi =

  2. [10]

    Inferring and Executing Programs for Visual Reasoning , booktitle =

    Justin Johnson and Bharath Hariharan and Laurens van der Maaten and Judy Hoffman and Li Fei. Inferring and Executing Programs for Visual Reasoning , booktitle =. 2017 , doi =

  3. [11]

    Hudson and Christopher D

    Drew A. Hudson and Christopher D. Manning , title =. 6th International Conference on Learning Representations,. 2018 , url =

  4. [12]

    Neural-Symbolic

    Kexin Yi and Jiajun Wu and Chuang Gan and Antonio Torralba and Pushmeet Kohli and Josh Tenenbaum , editor =. Neural-Symbolic. Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montr

  5. [13]

    Tenenbaum and Jiajun Wu , title =

    Jiayuan Mao and Chuang Gan and Pushmeet Kohli and Joshua B. Tenenbaum and Jiajun Wu , title =. 7th International Conference on Learning Representations,. 2019 , url =

  6. [14]

    DeepProbLog: Neural Probabilistic Logic Programming , booktitle =

    Robin Manhaeve and Sebastijan Dumancic and Angelika Kimmig and Thomas Demeester and Luc De Raedt , editor =. DeepProbLog: Neural Probabilistic Logic Programming , booktitle =

  7. [15]

    d'Avila Garcez , editor =

    Ivan Donadello and Luciano Serafini and Artur S. d'Avila Garcez , editor =. Logic Tensor Networks for Semantic Image Interpretation , booktitle =. 2017 , doi =

  8. [16]

    Neurosymbolic

    Artur d'Avila Garcez and Lu. Neurosymbolic. Artif. Intell. Rev. , volume =. 2023 , doi =

  9. [17]

    Coupling Large Language Models with Logic Programming for Robust and General Reasoning from Text , booktitle =

    Zhun Yang and Adam Ishay and Joohyung Lee , editor =. Coupling Large Language Models with Logic Programming for Robust and General Reasoning from Text , booktitle =. 2023 , doi =

  10. [18]

    Leveraging Large Language Models to Generate Answer Set Programs , booktitle =

    Adam Ishay and Zhun Yang and Joohyung Lee , editor =. Leveraging Large Language Models to Generate Answer Set Programs , booktitle =. 2023 , doi =

  11. [19]

    Proceedings of the 21st International Conference on Principles of Knowledge Representation and Reasoning (KR) , pages =

    Zhun Yang and Adam Ishay and Joohyung Lee , title =. Proceedings of the 21st International Conference on Principles of Knowledge Representation and Reasoning (KR) , pages =

  12. [20]

    Trinh and Yuhuai Wu and Quoc V

    Trieu H. Trinh and Yuhuai Wu and Quoc V. Le and He He and Thang Luong , title =. Nat. , volume =. 2024 , doi =

  13. [21]

    Theory and Practice of Logic Programming , year =

    Thomas Eiter and Nelson Higuera and Johannes Oetsch , title =. Theory and Practice of Logic Programming , year =

  14. [22]

    Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence,

    Thomas Eiter and Tobias Geibinger and Nelson Higuera and Johannes Oetsch , title =. Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence,. 2023 , doi =

  15. [23]

    Proceedings of the 17th International Workshop on Neural-Symbolic Learning and Reasoning (

    Thomas Eiter and Nelson Higuera Ruiz and Johannes Oetsch , title =. Proceedings of the 17th International Workshop on Neural-Symbolic Learning and Reasoning (. 2023 , url =

  16. [24]

    Thomas Eiter and Jan Hadl and Nelson Higuera Ruiz and Johannes Oetsch , title =. Proceedings of the 1st International Workshop on Next-Generation Language Models for Knowledge Representation and Reasoning (NeLaMKRR), co-located with the 21st International Conference on Princip...

  17. [26]

    Jimenez and John Yang and Alexander Wettig and Shunyu Yao and Kexin Pei and Ofir Press and Karthik R

    Carlos E. Jimenez and John Yang and Alexander Wettig and Shunyu Yao and Kexin Pei and Ofir Press and Karthik R. Narasimhan , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =

  18. [27]

    Jimenez and Alexander Wettig and Kilian Lieret and Shunyu Yao and Karthik Narasimhan and Ofir Press , editor =

    John Yang and Carlos E. Jimenez and Alexander Wettig and Kilian Lieret and Shunyu Yao and Karthik Narasimhan and Ofir Press , editor =. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering , booktitle =

  19. [28]

    CoRR , volume =

    Qingyun Wu and Gagan Bansal and Jieyu Zhang and Yiran Wu and Shaokun Zhang and Erkang Zhu and Beibin Li and Li Jiang and Xiaoyun Zhang and Chi Wang , title =. CoRR , volume =. 2023 , eprinttype =. 2308.08155 , doi =

  20. [29]

    Guanzhi Wang and Yuqi Xie and Yunfan Jiang and Ajay Mandlekar and Chaowei Xiao and Yuke Zhu and Linxi Fan and Anima Anandkumar , title =. Trans. Mach. Learn. Res. , volume =. 2024 , url =

  21. [30]

    Chi and Quoc V

    Jason Wei and Xuezhi Wang and Dale Schuurmans and Maarten Bosma and Brian Ichter and Fei Xia and Ed H. Chi and Quoc V. Le and Denny Zhou , editor =. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , booktitle =

  22. [31]

    Large Language Models are Zero-Shot Reasoners , booktitle =

    Takeshi Kojima and Shixiang Shane Gu and Machel Reid and Yutaka Matsuo and Yusuke Iwasawa , editor =. Large Language Models are Zero-Shot Reasoners , booktitle =

  23. [32]

    Le and Ed H

    Xuezhi Wang and Jason Wei and Dale Schuurmans and Quoc V. Le and Ed H. Chi and Sharan Narang and Aakanksha Chowdhery and Denny Zhou , title =. The Eleventh International Conference on Learning Representations,. 2023 , url =

  24. [33]

    Tree of Thoughts: Deliberate Problem Solving with Large Language Models , booktitle =

    Shunyu Yao and Dian Yu and Jeffrey Zhao and Izhak Shafran and Tom Griffiths and Yuan Cao and Karthik Narasimhan , editor =. Tree of Thoughts: Deliberate Problem Solving with Large Language Models , booktitle =

  25. [34]

    International Conference on Machine Learning,

    Luyu Gao and Aman Madaan and Shuyan Zhou and Uri Alon and Pengfei Liu and Yiming Yang and Jamie Callan and Graham Neubig , editor =. International Conference on Machine Learning,. 2023 , url =

  26. [35]

    Cohen , title =

    Wenhu Chen and Xueguang Ma and Xinyi Wang and William W. Cohen , title =. Trans. Mach. Learn. Res. , volume =. 2023 , url =

  27. [38]

    A Solver-in-the-Loop Framework for Improving LLMs on Answer Set Programming for Logic Puzzle Solving , booktitle =

    Timo Pierre Schrader and Lukas Lange and Tobias Kaminski and Simon Razniewski and Annemarie Friedrich , editor =. A Solver-in-the-Loop Framework for Improving LLMs on Answer Set Programming for Logic Puzzle Solving , booktitle =. 2026 , url =. doi:10.1609/AAAI.V40I30.39714 , t...

  28. [39]

    Theory and Practice of Logic Programming , year =

    Manuel Alejandro Borroto Santana and Katie Gallagher and Antonio Ielo and Irfan Kareem and Francesco Ricca and Alessandra Russo , title =. Theory and Practice of Logic Programming , year =

  29. [40]

    Neural module networks

    Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. Neural module networks. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016 , pages 39--48. IEEE Computer Society, 2016. doi:10.1109/CVPR.2016.12

  30. [41]

    The Claude 4 model family: Sonnet, opus, and haiku

    Anthropic . The Claude 4 model family: Sonnet, opus, and haiku. Anthropic technical report, 2025. URL https://www.anthropic.com/claude

  31. [42]

    Lawrence Zitnick, and Devi Parikh

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. VQA: visual question answering. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015 , pages 2425--2433. IE...

  32. [43]

    Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. Trans. Mach. Learn. Res., 2023, 2023. URL https://openreview.net/forum?id=YfZ4ZPt8zd

  33. [44]

    Fine-tuning llms for answer set programming

    Erica Coppolillo, Francesco Calimeri, Giuseppe Manco, Simona Perri, and Francesco Ricca. Fine-tuning llms for answer set programming. J. Intell. Inf. Syst., 64 0 (2): 0 653--685, 2026. doi:10.1007/S10844-025-01017-4. URL https://doi.org/10.1007/s10844-025-01017-4

  34. [45]

    Artur d'Avila Garcez and Lu \' s C. Lamb. Neurosymbolic AI: the 3rd wave. Artif. Intell. Rev., 56 0 (11): 0 12387--12406, 2023. doi:10.1007/s10462-023-10448-w

  35. [46]

    DeepSeek-V4 technical report

    DeepSeek-AI . DeepSeek-V4 technical report. DeepSeek-AI technical report, 2025. URL https://www.deepseek.com/

  36. [47]

    d'Avila Garcez

    Ivan Donadello, Luciano Serafini, and Artur S. d'Avila Garcez. Logic tensor networks for semantic image interpretation. In Carles Sierra, editor, Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI 2017, Melbourne, Australia, August...

  37. [48]

    A neuro-symbolic ASP pipeline for visual question answering

    Thomas Eiter, Nelson Higuera, and Johannes Oetsch. A neuro-symbolic ASP pipeline for visual question answering. Theory and Practice of Logic Programming, 2022

  38. [49]

    A logic-based approach to contrastive explainability for neurosymbolic visual question answering

    Thomas Eiter, Tobias Geibinger, Nelson Higuera, and Johannes Oetsch. A logic-based approach to contrastive explainability for neurosymbolic visual question answering. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, IJCAI 2023 , pa...

  39. [50]

    A modular neurosymbolic approach for visual graph question answering

    Thomas Eiter, Nelson Higuera Ruiz, and Johannes Oetsch. A modular neurosymbolic approach for visual graph question answering. In Proceedings of the 17th International Workshop on Neural-Symbolic Learning and Reasoning ( NeSy ) , CEUR Workshop Proceedings, pages 139--149. CEUR-...

  40. [51]

    Declarative knowledge distillation from large language models for visual question answering datasets

    Thomas Eiter, Jan Hadl, Nelson Higuera Ruiz, and Johannes Oetsch. Declarative knowledge distillation from large language models for visual question answering datasets. In Proceedings of the 1st International Workshop on Next-Generation Language Models for Knowledge Representat...

  41. [52]

    PAL: program-aided language models

    Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. PAL: program-aided language models. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, International Confer...

  42. [53]

    Multi-shot ASP solving with clingo

    Martin Gebser, Roland Kaminski, Benjamin Kaufmann, and Torsten Schaub. Multi-shot ASP solving with clingo. Theory Pract. Log. Program., 19 0 (1): 0 27--82, 2019. doi:10.1017/S1471068418000054

  43. [54]

    The stable model semantics for logic programming

    Michael Gelfond and Vladimir Lifschitz. The stable model semantics for logic programming. In Robert A. Kowalski and Kenneth A. Bowen, editors, Logic Programming, Proceedings of the Fifth International Conference and Symposium, Seattle, Washington, USA, August 15-19, 1988 (2 Vo...

  44. [55]

    Making the V in VQA matter: Elevating the role of image understanding in visual question answering

    Yash Goyal, Tejas Khot, Douglas Summers - Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, J...

  45. [56]

    Hinton, Oriol Vinyals, and Jeffrey Dean

    Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. Distilling the knowledge in a neural network. CoRR, abs/1503.02531, 2015. URL http://arxiv.org/abs/1503.02531

  46. [57]

    Hudson and Christopher D

    Drew A. Hudson and Christopher D. Manning. Compositional attention networks for machine reasoning. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings . OpenReview.net, 2018. URL ht...

  47. [58]

    Hudson and Christopher D

    Drew A. Hudson and Christopher D. Manning. GQA: A new dataset for real-world visual reasoning and compositional question answering. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019 , pages 6700--6709. Computer Visi...

  48. [59]

    Leveraging large language models to generate answer set programs

    Adam Ishay, Zhun Yang, and Joohyung Lee. Leveraging large language models to generate answer set programs. In Pierre Marquis, Tran Cao Son, and Gabriele Kern - Isberner, editors, Proceedings of the 20th International Conference on Principles of Knowledge Representation and Rea...

  49. [60]

    Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R

    Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R. Narasimhan. Swe-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7...

  50. [61]

    Lawrence Zitnick, and Ross B

    Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei - Fei, C. Lawrence Zitnick, and Ross B. Girshick. CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR ...

  51. [62]

    Lawrence Zitnick, and Ross B

    Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Judy Hoffman, Li Fei - Fei, C. Lawrence Zitnick, and Ross B. Girshick. Inferring and executing programs for visual reasoning. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29...

  52. [63]

    Large language models are zero-shot reasoners

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annua...

  53. [64]

    Shamma, Michael S

    Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li - Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei - Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations...

  54. [65]

    Deepproblog: Neural probabilistic logic programming

    Robin Manhaeve, Sebastijan Dumancic, Angelika Kimmig, Thomas Demeester, and Luc De Raedt. Deepproblog: Neural probabilistic logic programming. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol \` o Cesa - Bianchi, and Roman Garnett, editors, Advances in...

  55. [66]

    Tenenbaum, and Jiajun Wu

    Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B. Tenenbaum, and Jiajun Wu. The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, Ma...

  56. [67]

    GPT-5 system card

    OpenAI . GPT-5 system card. OpenAI technical report, 2025. URL https://openai.com/index/gpt-5-system-card/

  57. [68]

    Can llms solve ASP problems? insights from a benchmarking study

    Lin Ren, Guohui Xiao, Guilin Qi, Yishuai Geng, and Haohan Xue. Can llms solve ASP problems? insights from a benchmarking study. In Magdalena Ortiz, Renata Wassermann, and Torsten Schaub, editors, Proceedings of the 22nd International Conference on Principles of Knowledge Repre...

  58. [69]

    Question answering with LLMs and learning from answer sets

    Manuel Alejandro Borroto Santana, Katie Gallagher, Antonio Ielo, Irfan Kareem, Francesco Ricca, and Alessandra Russo. Question answering with LLMs and learning from answer sets. Theory and Practice of Logic Programming, 2025. doi:10.1017/S1471068425100343. URL https://doi.org/...

  59. [70]

    A solver-in-the-loop framework for improving llms on answer set programming for logic puzzle solving

    Timo Pierre Schrader, Lukas Lange, Tobias Kaminski, Simon Razniewski, and Annemarie Friedrich. A solver-in-the-loop framework for improving llms on answer set programming for logic puzzle solving. In Sven Koenig, Chad Jenkins, and Matthew E. Taylor, editors, Fortieth AAAI Conf...

  60. [71]

    Trinh, Yuhuai Wu, Quoc V

    Trieu H. Trinh, Yuhuai Wu, Quoc V. Le, He He, and Thang Luong. Solving olympiad geometry without human demonstrations. Nat., 625 0 (7995): 0 476--482, 2024. doi:10.1038/s41586-023-06747-5

  61. [72]

    Voyager: An open-ended embodied agent with large language models

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models. Trans. Mach. Learn. Res., 2024, 2024. URL https://openreview.net/forum?id=ehfRiF0R3a

  62. [73]

    Le, Ed H

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali,...

  63. [74]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, ...

  64. [75]

    Autogen: Enabling next-gen LLM applications via multi-agent conversation framework

    Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang. Autogen: Enabling next-gen LLM applications via multi-agent conversation framework. CoRR, abs/2308.08155, 2023. doi:10.48550/arXiv.2308.08155

  65. [76]

    Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press

    John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. ...

  66. [77]

    Coupling large language models with logic programming for robust and general reasoning from text

    Zhun Yang, Adam Ishay, and Joohyung Lee. Coupling large language models with logic programming for robust and general reasoning from text. In Anna Rogers, Jordan L. Boyd - Graber, and Naoaki Okazaki, editors, Findings of the Association for Computational Linguistics: ACL 2023,...

  67. [78]

    Learning to solve constraint satisfaction problems with large language models and answer set programming

    Zhun Yang, Adam Ishay, and Joohyung Lee. Learning to solve constraint satisfaction problems with large language models and answer set programming. In Proceedings of the 21st International Conference on Principles of Knowledge Representation and Reasoning (KR), pages 759--769, 2024 b

  68. [79]

    Tree of thoughts: Deliberate problem solving with large language models

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Adva...

  69. [80]

    Neural-symbolic VQA: disentangling reasoning from vision and language understanding

    Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum. Neural-symbolic VQA: disentangling reasoning from vision and language understanding. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol \` o Cesa - Bianchi, and Roman ...

  70. [81]

    Tenenbaum

    Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum. CLEVRER: collision events for video representation and reasoning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, ...

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.