Pith. sign in

REVIEW 5 cited by

NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.14890 v4 pith:PJRYCLCD submitted 2023-12-22 cs.AI cs.CCcs.CLcs.LG

classification cs.AIcs.CCcs.CLcs.LG
keywords llmsreasoningbenchmarkcomplexitynphardevalabilitiesabilitybenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Complex reasoning ability is one of the most important features of current LLMs, which has also been leveraged to play an integral role in complex decision-making tasks. Therefore, the investigation into the reasoning capabilities of Large Language Models (LLMs) is critical: numerous benchmarks have been established to assess the reasoning abilities of LLMs. However, current benchmarks are inadequate in offering a rigorous evaluation of the full extent of reasoning abilities that LLMs are capable of achieving. They are also prone to the risk of overfitting, as these benchmarks, being publicly accessible and static, allow models to potentially tailor their responses to specific benchmark metrics, thereby inflating their performance. Addressing these limitations, our research introduces a new benchmark, named NPHardEval. This benchmark is designed to evaluate the reasoning abilities of LLMs across a broad spectrum of 900 algorithmic questions, extending up to the NP-Hard complexity class. These questions are meticulously chosen to represent a wide range of complexity class below the NP-hard complexity class, offering a rigorous measure of the reasoning ability of LLMs. Through this study, we shed light on the current state of reasoning in LLMs, providing an objective and rigorous perspective through the comparison of LLMs' performance across complex classes. Moreover, this benchmark is designed with a dynamic update mechanism, where the datapoints are refreshed on a monthly basis. Such regular updates play a crucial role in mitigating the risk of LLMs overfitting to the benchmark, promoting a more accurate and reliable assessment of their reasoning capabilities. The benchmark dataset and code of NPHardEval are available at https://github.com/casmlab/NPHardEval.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

    cs.AI 2025-06 conditional novelty 6.0 of 10

    OPT-BENCH is a 30-task benchmark showing that LLM agents generally improve optimization results when they are given their own past solutions and feedback.

  2. LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LogiPlan introduces a three-task benchmark with controllable graph complexity, showing that although reasoning models excel at plan generation, all models degrade sharply on cycle detection and deep comparison questions.

  3. Integrating LLMs and Digital Twins for Adaptive Multi-Robot Task Allocation in Construction

    cs.RO 2025-06 conditional novelty 5.0 of 10

    The paper integrates digital twins, integer programming, and LLMs so that natural-language site updates can automatically adapt multi-robot construction task allocation.

  4. Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation

    cs.AI 2025-06 reject novelty 5.0 of 10

    A fixed-image, multi-task evaluation framework aims to detect data contamination in multimodal LLMs, but its judge is unvalidated and possibly self-referential, and the claimed harm to generalization is not supported ...

  5. From Long to Short: LLMs Excel at Trimming Own Reasoning Chains

    cs.AI 2025-09 conditional novelty 4.0 of 10

    EDIT searches across step-count prompts to find the shortest reasoning path the model consistently answers correctly.

Pith tools