Pith. sign in

REVIEW 1 cited by

Large Language Model for Science: A Study on P vs. NP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05689 v1 pith:QN4SSKLK submitted 2023-09-11 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsreasoningsciencelanguagelargeproblemproblemssocratic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In this work, we use large language models (LLMs) to augment and accelerate research on the P versus NP problem, one of the most important open problems in theoretical computer science and mathematics. Specifically, we propose Socratic reasoning, a general framework that promotes in-depth thinking with LLMs for complex problem-solving. Socratic reasoning encourages LLMs to recursively discover, solve, and integrate problems while facilitating self-evaluation and refinement. Our pilot study on the P vs. NP problem shows that GPT-4 successfully produces a proof schema and engages in rigorous reasoning throughout 97 dialogue turns, concluding "P $\neq$ NP", which is in alignment with (Xu and Zhou, 2023). The investigation uncovers novel insights within the extensive solution space of LLMs, shedding light on LLM for Science.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study

    cs.CL 2025-05 conditional novelty 5.0 of 10

    An LLM agent reproduces many known Russenorsk linguistic properties from a newly compiled dictionary, and can generate speculative Russenorsk translations, but the evaluation partly reflects prompt leakage.

Pith tools