Pith. sign in

REVIEW 4 cited by

DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.17167 v3 pith:VFS7P4RE submitted 2023-09-29 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords evaluationllmsdyvalbenchmarksdynamicreasoningsamplescomplexities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have achieved remarkable performance in various evaluation benchmarks. However, concerns are raised about potential data contamination in their considerable volume of training corpus. Moreover, the static nature and fixed complexity of current benchmarks may inadequately gauge the advancing capabilities of LLMs. In this paper, we introduce DyVal, a general and flexible protocol for dynamic evaluation of LLMs. Based on our framework, we build graph-informed DyVal by leveraging the structural advantage of directed acyclic graphs to dynamically generate evaluation samples with controllable complexities. DyVal generates challenging evaluation sets on reasoning tasks including mathematics, logical reasoning, and algorithm problems. We evaluate various LLMs ranging from Flan-T5-large to GPT-3.5-Turbo and GPT-4. Experiments show that LLMs perform worse in DyVal-generated evaluation samples with different complexities, highlighting the significance of dynamic evaluation. We also analyze the failure cases and results of different prompting methods. Moreover, DyVal-generated samples are not only evaluation sets, but also helpful data for fine-tuning to improve the performance of LLMs on existing benchmarks. We hope that DyVal can shed light on future evaluation research of LLMs. Code is available at: https://github.com/microsoft/promptbench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Automated Capability Discovery via Foundation Model Self-Exploration

    cs.LG 2025-02 conditional novelty 6.0 of 10

    ACD automatically generates thousands of open-ended tasks and clusters them into dozens of capability and failure categories, with LLM-vs-human scoring agreement (F1 = 0.86).

  2. Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation

    cs.AI 2025-06 reject novelty 5.0 of 10

    A fixed-image, multi-task evaluation framework aims to detect data contamination in multimodal LLMs, but its judge is unvalidated and possibly self-referential, and the claimed harm to generalization is not supported ...

  3. Unbiased Evaluation of Large Language Models from a Causal Perspective

    cs.AI 2025-02 reject novelty 5.0 of 10

    The paper argues that perturbing benchmark questions with rule-based interventions gives a less contaminated, more interpretable evaluation of LLMs than static benchmarks or agent-generated questions.

  4. Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era

    cs.LG 2025-08 unverdicted novelty 1.0 of 10

    A tutorial proposal outlining how generative models can synthesize data across modalities for data mining, with no new research results.

Pith tools